Staff Site Reliability Engineer (x/f/m)
DoctolibParisFull TimeEngineering
Remotely
awsdatadogkubernetesgcpterraformargocdprometheusopentelemetry
Job Description
📋 Description
- Lead large-scale cross-cutting reliability initiatives across the platform, spanning infrastructure
- Identify and drive improvements to incident detection, response, and postmortem analysis
- Define and evolve SLOs, error budgets, and alerting standards across multiple product teams
- Take part in the on-call rotation, and actively contribute to improving our on-call experience by
- Serve as a mentor and technical coach to senior engineers, helping elevate the craft of reliability
- Influence strategic decisions by providing technical guidance to leadership and representing
🎯 Requirements
- 8+ years in SRE, platform engineering, or infrastructure roles within a large-scale, multi-team
- Proven experience with cloud platforms such as AWS, GCP, or Azure
- Strong experience with containerization and orchestration technologies, Kubernetes is a must
- Implemented and operated SLIs, SLOs, and error budgets in production
- Experience managing on-call rotations and leading incident response in high-stakes environments
- Strong systems engineering background with fluency in at least one backend language (Go, Python
🎁 Benefits
- Tech stack: Kubernetes, Terraform, AWS / GCP, Prometheus, OpenTelemetry, Datadog, ArgoCD
- Hybrid work setup (up to 2 remote days per week)
- Relocation support in case of international mobility
- Access to AI tools and dedicated training
- Other Doctolib benefits listed in job details