Job Description
📋 Description Own production reliability across the full stack. Scale the on-call culture with robust runbooks and processes. Instrument monitoring, alerts, and observability to surface issues early. Drive the SRE roadmap: define investments, SLOs, chaos engineering. Make engineers faster with tooling, feedback loops, smoother deployments. Collaborate across squads with backend, ML, and frontend to embed reliability. 🎯 Requirements Senior-level experience in SRE, infra, or backend engineering. Hands-on cloud infra (GCP preferred; AWS/Azure OK to learn a new stack). Structured on-call: alerting policies, runbooks, post-mortems. Strong IaC fundamentals; Terraform; reviewing PRs. Autonomy: self-prioritize and drive to completion. Clear communication across distributed teams, including during incidents. Bonus: Kubernetes, GCP tooling (Cloud Run, GKE, Pub/Sub), PostgreSQL at scale. 🎁 Benefits Compensation and Equity: Competitive salary and stock options. Comprehensive health plans: 100% coverage for Medical, Dental, Vision. Time Off: Unlimited PTO and 11 national holidays. Health Comes First: Unlimited sick leave. Parental Leave: Paid leave for new parents. Trust & accountability: Full ownership of your time and schedule.