Remotely
awskubernetesterraformgithub actionsprometheusgrafanaopentelemetrygitlab ci
Job Description
📋 Description
- Lead discovery, design, delivery of reliability infra initiatives
- Contribute to platform architecture and roadmaps
- Define SLOs, SLIs, error budgets, and observability standards
- Use incident metrics to drive reliability strategy
- Scale production Kubernetes and container infra
- Build and manage cloud infra with AWS or similar
🎯 Requirements
- Solid SRE/DevOps/Platform Engineering experience
- Production Kubernetes with Docker and container ecosystem
- Production cloud infra with AWS or comparable cloud
- Terraform and IaC principles
- Observability with OpenTelemetry, Grafana, Prometheus
- CI/CD pipelines with GitLab CI, GitHub Actions
🎁 Benefits
- 100% remote work, asynchronous environment
- Flexible hours and async-friendly culture
- Flexible paid time off and 16 weeks parental leave
- Home office budget, learning and wellness budget
- Mental health support and stock options
- Competitive, location-aware compensation; global market parity
Back to all jobs