Senior Site Reliability Engineer
2KRemotely
kubernetesterraformgitopsgkeargocdprometheusekspulumi
Job Description
📋 Description
- Hands-on technical leader shaping production infrastructure across multi-clouds and regions.
- Own platform reliability from architecture review through production operations.
- Collaborate with network engineers, systems architects, and game studios.
🎯 Requirements
- 5+ years in SRE, platform engineering, or equivalent infra at production scale.
- Deep Kubernetes experience in cloud environments (EKS/GKE).
- Strong IaC with Terraform/Pulumi; hands-on with Helm, Terragrunt, GitOps (ArgoCD/GitHub Actions).
- Observability stack experience: Datadog, Prometheus, Grafana; SLI/SLO and error budget fluency.
- Production-quality code in Go, Python, or TypeScript; Linux, TCP/IP, DNS, TLS fundamentals.
- Incident response leadership and post-mortem ownership.
🎁 Benefits
- Great Company Culture; growth opportunities and employee wellbeing programs.
- Comprehensive benefits including health coverage and life insurance.
- Perks such as gym reimbursements and learning platforms.