North AmericaFull TimeEngineering
Remotely
awscloudkubernetessreincident managementobservabilityslodre
Job Description
📋 Description
- Define and execute Everbridge's global production operations and reliability strategy.
- Lead and develop global SRE and DRE teams for platform reliability and operational engineering.
- Own the operational excellence of AWS and Kubernetes-based cloud platform: scalability, resilience
- Establish SLOs, error budgets, production readiness, capacity planning, observability, and
- Drive incident management, change governance, release management, and post-incident reviews.
- Lead disaster recovery, resilience testing, and business continuity across global environments.
🎯 Requirements
- 15+ years in production operations, cloud infrastructure, platform engineering, SRE, or related
- Experience leading global production operations for large-scale SaaS/cloud platforms.
- Deep expertise in AWS, Kubernetes, cloud-native architectures, distributed systems
- Strong SRE principles, incident management, observability, DR, change management, CI/CD, IaC.
- Proven track record building and leading global engineering/operations teams.
- Excellent executive communication and cross-functional leadership.
🎁 Benefits
- Executive-level scope with strategic impact across global production operations.
- Competitive compensation and comprehensive benefits package.
- Collaborative environment partnering with Engineering, Product, Security, Support, and Success
- Opportunity to influence reliability culture and cloud-native practices across the company.