Remotely
awsdatadogmysqlprometheusgrafanapmmaurora mysqlperformance insights
Job Description
📋 Description
- Own the reliability, performance, and availability of all production databases (Aurora MySQL on
- Manage database alerting end to end – observability pipeline to Datadog; fold after-hours alerts
- Educate engineering teams about proper database hygiene and best-performing queries.
- Implement automated safeguards, including automatic termination of long-running queries and visible
- Author and maintain database runbooks for quick, consistent incident response.
- Diagnose, optimize, and re-index problematic queries to reduce latency and load.
🎯 Requirements
- 6+ years in database reliability engineering, administration, or engineering with production
- Deep MySQL expertise — query optimization, execution plans, indexing, replication; Aurora MySQL
- Experience operating large multi-reader Aurora/RDS clusters at terabyte scale with read-routing.
- Strong observability tooling knowledge (Datadog, PMM, Performance Insights, Prometheus/Grafana).
- Automation scripting (Python, Bash) and infrastructure-as-code (Terraform).
- Solid AWS RDS/Aurora experience and cost optimization skills.
🎁 Benefits
- Remote/hybrid environment
- Competitive salaries
- Potential equity compensation for outstanding performance
- Flexible PTO
- Company-paid disability and life insurance
- Medical, dental, and vision insurance
Back to all jobs