Site Reliability Engineer
XsollaBakuFull TimeEngineering
Remotely
datadogkubernetesterraformgkeci/cdopentelemetrygitlab cihelm
Job Description
📋 Description
- Own the application-level infrastructure: Helm charts, Terraform configurations, Kubernetes
- Own the domain's observability: design and implement SLOs/SLIs, monitors, alerts, and dashboards
- Help to set up and evolve CI/CD pipelines for domain services (GitLab CI, GitHub Actions)
- Perform capacity planning and performance tuning ahead of expected load - product launches, sales
- Run Production Readiness Reviews for new services and major changes; define and enforce what
- Support domain incident response: assist with deep investigation of complex incidents, contribute
🎯 Requirements
- 3+ years of proven SRE, DevOps, or platform engineering experience: on-call or incident response
- Software development background: built and shipped backend services and written production-quality
- Hands-on Kubernetes experience: Helm, manifests, deploy strategies, debugging production
- Strong observability practice: monitors, dashboards, SLOs/SLIs on Datadog or Prometheus/Grafana
- Terraform/Terragrunt for IaC; GCP experience (IAM, networking, managed services)
- Experience building and maintaining CI/CD pipelines (GitLab CI and/or GitHub Actions) and scripting
🎁 Benefits
- Comprehensive Benefits Program: medical, dental, vision, PTO, and career roadmap
- Professional development opportunities and ongoing training