Job Description
📋 Description Own Data Replication infra: Kubernetes, CI/CD, secrets, networking, cloud config. Partner with product engineers to reliably integrate features with infrastructure. Maintain observability, alerting, and anomaly detection with AI/LLM focus. Maintain AI-augmented release tooling: canaries, progressive rollouts, rollback. Build self-serve tooling, write runbooks, and coach engineers to own more of their stack. 🎯 Requirements 7+ years in infrastructure, platform engineering, SRE, or DevOps. Hands-on ownership of Kubernetes, Helm, and Terraform in production environments. Experience with observability stacks (Prometheus, Grafana, Datadog) and on-call operations. Experience with CI/CD pipeline ownership and developer tooling. Ability and willingness to read backend code to understand and instrument failures. Fluency with AI tools - LLMs and agentic frameworks to automate and reduce toil. Startup-ready mindset: comfortable with ambiguity, moving fast, owning problems end-to-end. 🎁 Benefits Flexible PTO with at least 25 days off annually. 16 weeks fully paid parental leave for all parents. Comprehensive medical, dental, and vision coverage for employees and dependents. 401(k) retirement plan. Professional development budget, conference sponsorship, and book reimbursement. Commuter benefits and monthly internet reimbursement. Breakfast and lunch in our San Francisco office. Collaborative, in-person culture focused on learning, growth, and impact.