Remotely
awskubernetesterraformgithub actionsprometheusgrafanaopentelemetrygitlab ci
Job Description
📋 Description
- Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives
- Contribute to platform architecture, infrastructure tooling, technical roadmaps, and engineering
- Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies
- Use operational and incident metrics to identify systemic reliability issues and influence the
- Resolve cross-team infrastructure and platform requests while turning recurring problems into
- Operate and scale production Kubernetes environments and associated container infrastructure.
🎯 Requirements
- Solid professional experience in Site Reliability Engineering, DevOps, Platform Engineering, or a
- Strong hands-on experience operating and scaling Kubernetes in production, including Docker and the
- Proven experience designing, building, and managing production cloud infrastructure using AWS or a
- Strong practical expertise with Terraform and infrastructure-as-code principles.
- Hands-on experience with reliability engineering frameworks, including SLOs, SLIs, error budgets
- Strong observability experience with technologies such as OpenTelemetry, Grafana, Prometheus, or
🎁 Benefits
- 100% remote work, with the ability to work from anywhere.
- Flexible working hours within an async-first working environment.
- Flexible paid time off to support a healthy balance between work and personal life.
- 16 weeks of paid parental leave.
- Budget for co-working spaces, learning, and wellness, including gym memberships.
- Mental health support services.
Back to all jobs