Sr Platform Monitoring Engineer
DatabricksNorth AmericaFull TimeEngineering
Remotely
pythonawsdockerkubernetesprometheusgrafanaelkpagerduty
Job Description
📋 Description
- Lead platform incident investigations and cross-functional coordination
- Design observability solutions and alerting workflows
- Drive reliability improvements across Databricks Platform
- Mentor engineers on observability patterns and metrics
- Participate in on-call rotation
🎯 Requirements
- Minimum of 6 years in SRE/DevOps/Production Eng or similar
- Production experience with AWS/Azure/GCP and Kubernetes
- Hands-on with ELK, Prometheus, Grafana, PagerDuty
- Strong Python or similar scripting skills for automation
- Experience with end-to-end incident lifecycle and post-mortems
- BS/MS/PhD in CS/CE or related field
🎁 Benefits
- Benefits and perks details available per region
- Inclusive, diverse culture
Back to all jobs