Senior Site Reliability Engineer I
BrazeVancouverFull TimeEngineering
Remotely
gopythonkubernetesansibleterraformrubynginxingress
Job Description
📋 Description
- Maintain internal services and platforms to maximize uptime and reliability
- Work on ingress fleets and API ingestion layers for scalable, real-time traffic
- Collaborate with product teams to translate requirements into resilient architectures
- Lead NGINX and Kubernetes ingress infrastructure and scaling routines
- Participate in on-call rotations and blameless postmortems to drive improvements
- Design systems with SLIs/SLOs and capacity planning for enterprise-grade SLAs
🎯 Requirements
- 5+ years in DevOps or SRE in a high-scale prod environment
- Deep expertise with NGINX and Kubernetes
- Proficiency in Ruby and/or Go (or Python/Java) for automation
- Infrastructure as Code experience (Terraform, Ansible, Chef or similar)
- Strong Linux/Unix fundamentals and TCP/IP networking
- Systems thinking and collaborative, async teamwork in a global engineering context
🎁 Benefits
- Competitive compensation with equity
- Retirement and equity plans; flexible PTO
- Comprehensive medical/dental/vision and disability coverage
- Family services, learning stipend, and opportunities for community involvement
- Great Place to Work culture and supportive resources
Back to all jobs