Remotely
awskubernetesterraformgithub actionsprometheusgrafanaopentelemetrygitlab ci
Job Description
📋 Description
- Lead discovery, design, and delivery of reliability and infra initiatives.
- Shape platform architecture, tooling, and roadmaps to boost reliability and developer experience.
- Define reliability practices: SLOs, SLIs, error budgets, alerts, observability.
- Use metrics to identify systemic reliability issues and drive strategy.
- Resolve cross-team infra requests; automate recurring problems; write runbooks.
- Operate and scale production Kubernetes environments and container infra.
🎯 Requirements
- Solid experience in Site Reliability Engineering/DevOps/Platform Eng or related.
- Hands-on Kubernetes in production (Docker and container ecosystem).
- Production cloud infra with AWS or comparable provider.
- Terraform and IaC principles experience.
- Experience with SLOs/SLIs, error budgets, alerting, incident management.
- Observability with OpenTelemetry, Grafana, Prometheus or similar.
🎁 Benefits
- 100% remote work with location flexibility.
- Async-first environment and flexible hours.
- Flexible paid time off for work-life balance.
- 16 weeks paid parental leave.
- Budget for coworking spaces, learning, wellness; gym memberships.
- Mental health support services.
Back to all jobs