Engineering Lead - Site Reliability Engineer
Heidi HealthRemotely
cloudkubernetesterraformobservabilityiacslos
Job Description
📋 Description
- Lead and grow the platform team covering cloud infrastructure and reliability. Set direction, hire
- Own Heidi's multi-region cloud footprint end to end: compute, networking, data stores, Kubernetes
- Set reliability bar with SLIs/SLOs; run error budgets and influence roadmap decisions.
- Lead incident response; define severity, escalation, comms, and blameless post-incident reviews.
- Partner with Release Manager to sharpen deployment processes (flags, canary, staged rollouts
- Collaborate with product teams to ensure platform enables faster engineering across regions.
🎯 Requirements
- Significant experience running cloud infrastructure or SRE teams at scale; regulated environments
- Hands-on with major cloud provider, Kubernetes, Terraform/IaC, and modern observability stacks.
- Proven reliability practice: defined SLOs, enforced error budgets, incidents, post-incident reviews.
- Track record designing for high availability across regions, with DR/testing and data residency
- Service mindset: measure platform impact by enabling other teams to move faster.
🎁 Benefits
- Learning and development budget of $1,000 annually.
- $150/month health and wellness allowance.
- $500 home office budget.
- 26 weeks primary parental leave and 18 weeks secondary.
- Fertility support up to $10,000.
- Four weeks work from anywhere per year.