Remotely
awskubernetesgolangterraformgithub actionsprometheusopentelemetrygitlab ci
Job Description
📋 Description
- Lead discovery, design, delivery of reliability initiatives
- Contribute to platform architecture and tooling
- Define reliability practices: SLOs, SLIs, alerts
- Use metrics to identify issues and influence strategy
- Scale Kubernetes environments and container infra
- Develop cloud infra with AWS or comparable platforms
🎯 Requirements
- Experience in SRE/DevOps/Platform Eng or related
- Hands-on Kubernetes in production (Docker, containers)
- Production cloud infra with AWS or equivalent
- Terraform and IaC proficiency
- Experience with SLOs/SLIs, error budgets, incident mgmt
- Observability tech: OpenTelemetry, Grafana, Prometheus
🎁 Benefits
- 100% remote work from anywhere
- Async-first working environment
- Flexible paid time off
- 16 weeks paid parental leave
- Budget for coworking, learning, wellness
- Mental health support services
Back to all jobs