Also available on
Job Description
📋 Description Lead and grow the SRE team; set on-call cadence and capacity planning. Define SLOs/SLIs and error budgets; embed in release processes. Lead major incident responses; conduct blameless post-incident reviews. Champion proactive reliability via chaos engineering and game days. Build AI-native operations: anomaly detection and automated triage. Improve observability and platform reliability; manage on-call escalation. 🎯 Requirements 10+ years in SRE/DevOps/Platform Eng. 4+ years direct people management (remote/on-call teams). Ownership of reliability outcomes; define SLOs/SLIs/error budgets. Deep Azure experience (3+ yrs) with AKS, networking, identity; AWS/GCP ok. Incident command experience; lead severity-1 incidents; blameless reviews. Bachelor's degree in CS/IT/Engineering or equivalent. 🎁 Benefits Free premium health, dental, life, and vision insurance. Generous 401(k) match. Paid sick leave; accrual policy per applicable law. Company events, virtual happy hours and team-building activities. Unlimited PTO and flexible vacation. Virtual yoga, meditation or boot camp classes offered daily.