Senior Site Reliability Engineer
MegaportSao PauloFull TimeEngineering
Remotely
kubernetesansibleterraformprometheusbashgrafanaelkloki
Job Description
📋 Description
- Continuously improve Latitude.sh’s platform reliability and performance
- Design, build, and maintain tools to automate operational tasks and incident response
- Implement and improve observability solutions, including monitoring, alerting, and tracing
- Collaborate with engineering and platform teams to design scalable and resilient systems
- Participate in on-call rotations and lead post-incident reviews focusing on learning
- Develop and document processes and runbooks for operational excellence
🎯 Requirements
- Strong verbal and written English communication skills
- Advanced knowledge of Linux/Unix systems in production environments
- Experience with Kubernetes and container orchestration
- Proficiency with infrastructure automation tools (e.g., Terraform, Ansible)
- Experience with observability stacks (Prometheus, Grafana, Loki, ELK)
- Familiarity with scripting and programming languages such as Bash, Python, Go, or Ruby
🎁 Benefits
- Contractor (PJ)
- Paid Time Off
- Competitive Compensation
- Wellhub (former Gympass)
- Annual Bonus based on company and team performance
- Flexible work hours