Senior Site Reliability Engineer
CertifyOSJob Description
📋 Description Own the end-to-end operational lifecycle and deployment workflows. Design for reliability; automate infrastructure for production readiness. Build and maintain IaC, CI/CD, and observability tooling. Lead incident response, root-cause analysis, and postmortems. Improve observability and reliability at scale with SLIs/SLOs and data signals. Collaborate across teams and mentor others on reliability practices. 🎯 Requirements 5+ years in SRE/DevOps/Platform/Infra at scale. Linux administration, incident response, root-cause analysis; mentor teams. GCP experience: GKE, Cloud Run, containerized workloads. Terraform and/or Pulumi for IaC; CI/CD pipelines experience. Deployment patterns: rolling, blue/green, canary; rollback strategies. Observability: Golden Signals; design SLIs/SLOs; dashboards; Prometheus/Grafana experience. 🎁 Benefits 100% coverage of health, dental, vision for US employees. Unlimited PTO with at least two weeks off annually (US). Inclusive, equal opportunity employer; welcoming all backgrounds. Pay transparency and accommodations during the application process.