Remotely
awskubernetesterraformci/cdgithub actionsprometheusgrafanaopentelemetry
Job Description
📋 Description
- Lead discovery, design and delivery of reliability initiatives for Kubernetes and cloud infra.
- Shape platform architecture and roadmaps; improve reliability and developer experience.
- Define SLOs, SLIs, error budgets, alerts, and observability standards.
- Operate and scale production Kubernetes and cloud infrastructure (AWS or comparable).
- Develop infrastructure as code with Terraform and CI/CD automation.
- Mentor engineers and collaborate asynchronously across a global team.
🎯 Requirements
- Senior SRE/DevOps/Platform with hands-on Kubernetes, Docker, and container ecosystem.
- Production cloud infra experience (AWS or equivalent).
- Terraform and IaC principles; observability with OpenTelemetry, Grafana, Prometheus.
- CI/CD design using GitLab/GitHub Actions or similar; Golang/Bash scripting.
- Experience with AI/agentic workflows in infra or dev workflows a plus.
- Strong communication in asynchronous, globally distributed environments.
🎁 Benefits
- 100% remote work with asynchronous-first environment.
- Flexible hours and paid time off; 16 weeks parental leave.
- Home office budget, wellness stipend, stock options.
- Competitive, location-aware USD salary; autonomy in global org.
Back to all jobs