SRE Engineer
JobgetherRemotely
awsdockerkubernetesazuregcpobservabilityci/cdapm
Job Description
📋 Description
- Strengthen reliability, resilience, and performance of digital environments.
- Work across cloud infrastructure, Kubernetes, observability, automation, and incident management to
- Define and monitor reliability metrics to drive measurable improvements.
- Collaborate with multidisciplinary teams to embed reliability and observability from design onward.
- Automate operational toil and improve scalability through IaC and capacity planning.
🎯 Requirements
- Site Reliability Engineer (SRE) or equivalent reliability/DevOps/infrastructure role experience.
- Cloud environments especially GCP, AWS, and/or Azure.
- Kubernetes and Docker proficiency.
- Observability, monitoring, alerting, and APM solutions experience.
- Strong understanding of SRE concepts: SLI, SLO, SLA, MTTR, MTTD, and error budgets.
- Production incidents management and strong troubleshooting skills.
🎁 Benefits
- Meal allowance (Vale Refeição).
- Food allowance (Vale Alimentação).
- Home office allowance, medical, dental, and life insurance.
- Birthday day off and wellness programs.
- Structured onboarding and access to continuous learning.
- Internal initiatives for knowledge sharing and professional development.
Back to all jobs