SRE Engineer
JobgetherRemotely
dockerkubernetesansibleterraformobservabilityci/cdprometheusgrafana
Job Description
📋 Description
- Define, monitor, and improve reliability indicators (SLIs, SLOs, SLAs, MTTR, MTTD, error budgets).
- Implement observability, monitoring, alerting, and APM across apps and infra.
- Monitor latency, traffic, errors, saturation, availability, performance.
- Prevent, investigate, resolve incidents; minimize impact to users/business.
- Root cause analyses; define corrective actions to avoid recurrence.
- Identify risks, bottlenecks, single points of failure; strengthen resilience.
🎯 Requirements
- Proven SRE/Site Reliability Engineer experience or equivalent.
- Cloud experience with GCP, AWS, or Azure.
- Hands-on Kubernetes and Docker.
- Experience with observability, monitoring, alerting, and APM.
- Strong SRE concepts/metrics (SLI, SLO, SLA, MTTR, MTTD, error budgets).
- Experience managing incidents; troubleshooting production environments.
🎁 Benefits
- Meal and food allowance.
- Home office allowance.
- Medical, dental, life insurance.
- Birthday Day Off.
- TotalPass / Wellhub access.
- Wellness support and partner discounts.
Back to all jobs