Staff Site Reliability Engineer (x/f/m)
DoctolibRemotely
awsdatadogkubernetesgcpterraformargocdprometheusopentelemetry
Job Description
📋 Description
- Lead large-scale reliability initiatives across the platform
- Drive incident detection, response, and postmortems
- Define SLOs, error budgets, and alert standards
- Participate in on-call rotations and improve on-call experience
- Mentor engineers and guide reliability across teams
- Partner with software teams to embed reliability early in development
🎯 Requirements
- 8+ years SRE/platform/infrastructure in large-scale env
- Cloud experience with AWS, GCP, or Azure
- Kubernetes expertise; deployment and scaling strategies
- Experience operating SLI/SLOs and error budgets
- On-call leadership and incident response experience
- Strong systems background with Go/Python/Ruby
🎁 Benefits
- Hybrid work setup (up to 2 remote days per week)
- Competitive compensation and health benefits
- Relocation support for international mobility
- Access to AI tools and dedicated training
- Flexible working policies and supportive environment
- Germany-specific perks like Deutschlandticket