Senior Manager, Site Reliability Engineering (SRE)
NIUMRemotely
awsdatadogkubernetesterraformopensearchprometheusgrafana
Job Description
📋 Description
- Lead, mentor, and grow a team of SREs across multiple time zones.
- Own reliability roadmap, SLIs/SLOs, and error budgets for critical services.
- Drive incident management, on-call structure, and blameless postmortems.
- Collaborate with product, security, and infra teams to ensure always-on reliability.
- Build scalability and operational readiness into the software lifecycle.
- Champion automation to reduce toil and improve MTTD/MTTR.
🎯 Requirements
- 10+ years of software/infra/SRE experience; 4+ years in people management or tech leadership
- Proven track record scaling high-availability, transaction-heavy systems
- Deep expertise in cloud infrastructure (AWS), Kubernetes, container orchestration, and IaC
- Strong observability stack experience (Prometheus, Grafana, Datadog, ELK/OpenSearch) and alerting
- Experience with SRE governance: SLOs, error budgets, capacity planning, production readiness
- Excellent communication skills; ability to present reliability concepts to peers and executives.
🎁 Benefits
- Hybrid work environment with 3 days/week in the office.
- Competitive compensation, performance bonuses, equity for specific roles, and recognition programs.
- Wellbeing coverage and employee assistance program; generous vacation; learning stipend.
- Global exposure with 100+ markets and diverse, multinational teams.
Back to all jobs