Remotely
awskubernetesterraformgithub actionsprometheusgrafanaopentelemetrygitlab ci
Job Description
📋 Description
- Lead discovery, design, and delivery of reliability initiatives
- Shape platform architecture and tooling priorities for reliability
- Define SLOs, SLIs, error budgets, alerts, observability
- Use metrics to identify systemic reliability issues
- Resolve cross-team requests, automate and document recurring problems
- Operate and scale production Kubernetes and container infra
🎯 Requirements
- Experience in SRE/DevOps/Platform Engineering
- Production Kubernetes with Docker and container ecosystem
- Production cloud infra using AWS or similar
- Terraform and IaC expertise
- Experience with SLOs/SLIs, alerting, incident mgmt
- Observability with OpenTelemetry, Grafana, Prometheus
🎁 Benefits
- 100% remote work, anywhere
- Async-friendly hours
- Flexible paid time off
- 16 weeks parental leave
- Budget for coworking spaces, learning, wellness
- Mental health support services
Back to all jobs