Remotely
awskubernetesterraformci/cdprometheusgrafanaopentelemetrygitlab ci
Job Description
📋 Description
- Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives
- Contribute to platform architecture, infrastructure tooling, roadmaps, and engineering priorities
- Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies
- Use operational and incident metrics to identify systemic reliability issues and influence
- Operate and scale production Kubernetes environments and container infrastructure.
- Build and manage cloud infrastructure with AWS or comparable platforms with emphasis on reliability
🎯 Requirements
- Solid experience in Site Reliability Engineering/DevOps/Platform Engineering or related discipline.
- Hands-on production Kubernetes experience (Docker and container ecosystem).
- Production cloud infrastructure experience with AWS or equivalent.
- Terraform and IaC principles expertise.
- Experience with SLOs/SLIs, error budgets, alerting, and incident management.
- Observability experience with OpenTelemetry, Grafana, Prometheus or similar.
🎁 Benefits
- 100% remote work, asynchronous-first environment.
- Flexible hours and async-friendly schedule.
- Flexible PTO, 16 weeks of parental leave, home office budget, and IT equipment.
- Mental health support, gym/wellness budget, and stock options.
- Competitive, location-aware USD salary range ($53,300 – $119,850) and global pay considerations.
- Opportunity to influence platform architecture and reliability practices in a distributed team.
Back to all jobs