Senior Site Reliability Engineer
AlpacaAlso available on
Job Description
📋 Description Operate production day-to-day: oncall, incident response, postmortems, follow-ups. Own reliability practice: define/refine SLIs/SLOs and error budgets. Strengthen observability across metrics, logs, traces, alerts. Ship infrastructure through code in a GitOps workflow: cloud resources & Kubernetes workloads. Look after PostgreSQL: performance tuning, online migrations, HA/DR, CDC pipelines. Mentor engineers on reliability and database fundamentals through code and design reviews. 🎯 Requirements 4+ years in SRE, DevOps, Platform/Infrastructure, or backend engineering with production ownership. Hands-on experience operating production services on Kubernetes; shipping infrastructure as code in Solid working knowledge of PostgreSQL in production. Cloud networking fundamentals (VPCs, routing, load balancing, DNS, TLS). Comfortable with a modern observability stack; proficient with Linux at the operator level. Incident response experience with calm, structured debugging and postmortems driving change. 🎁 Benefits Competitive Salary & Stock Options Health Benefits New Hire Home-Office Setup: One-time USD $500 Monthly Stipend: USD $150 per month via a Brex Card