Remotely
automationcloudincident responsesremonitoringci/cd
Job Description
📋 Description
- Lead, coach, mentor, and develop a high-performing SRE team, establishing ownership and continuous
- Drive reliability practices across cloud services: SLOs, SLIs, monitoring, alerting, capacity
- Partner with engineering/platform teams to improve architecture, scalability, resilience
- Own incident management: major incident coordination, post-incident reviews, root-cause analysis
- Champion automation, tooling, and engineering practices reducing manual toil and increasing
- Analyze operational data and service metrics to identify reliability gaps and communicate progress
🎯 Requirements
- 5+ years in SRE, DevOps, infrastructure, cloud operations, or related ops discipline, with
- Strong understanding of cloud platforms, distributed systems, production operations, and SRE
- Hands-on with SLOs/SLIs, monitoring, alerting, incident response, and post-incident reviews.
- Proven ability to hire, mentor, coach, and develop engineers; foster inclusive, accountable culture.
- Strong communication to explain technical concepts to engineers, product teams, and leaders.
- Experience with infrastructure automation, CI/CD, disaster recovery, cost optimization, or
🎁 Benefits
- Anticipated salary range of $151,000–$200,000, with actual compensation based on skills, location
- Equity in non-qualifying stock options.
- High-quality health benefits.
- Retirement plan with employer matching.
- Flexible Time Off and Paid Time Off benefits.
- Career development and growth opportunities; remote-friendly environment.