San FranciscoFull TimeEngineering
Remotely
awskubernetesterraformobservabilityargocdci
Job Description
📋 Description
- Own infra and data platform roadmap; scale for more traffic; target 99.95% uptime; manage cloud
- Shape agentic production operations; enable engineers to delegate routine work to agents
- Define reliability practice with SLI/SLOs; manage incident response and post-mortems
- Improve developer experience via CI standardization and self-service tooling; enable IaC without
- Lead incident command during outages; turn fixes into systemic improvements
- Own vendor strategy and infra budget; negotiate contracts and migrate vendors as needed
🎯 Requirements
- 8+ years infrastructure/platform engineering; 4+ years leading/managing engineers at high-growth
- AI-driven ops background; can name the system designed and impact on team work
- Decisive with incomplete info; ship, measure, adjust based on data
- Managed managers or tech leads; scaled teams while preserving culture/performance
- Deep hands-on AWS, Kubernetes, Terraform, ArgoCD; modern observability tooling
- Strong business acumen: cloud spend, velocity, reliability as enablement
🎁 Benefits
- Flexible PTO
- In-office lunch twice per week for NYC and SF, plus late-night dinner stipends
- Wellhub corporate wellness platform with gym discounts and mental health resources
- Annual professional development stipend
- $5,000 bonus at 5 years for recharging
- 401k with company contributions from day 1
Back to all jobs