Job Description
📋 Description Design, implement, maintain, and monitor reliable production systems at scale. Lead incident response, mitigate production issues, and conduct post mortem analysis. Proactively monitor performance, analyze system failures, identify bottlenecks, and propose Create and support observability/monitoring tools and vendor integrations. Drive the growth of a reliability culture, promoting cross-functional collaboration towards Train and mentor other engineers. 🎯 Requirements 5+ years of experience as a reliability-focused engineer in a fast-paced, rapidly growing Deep understanding of tooling and application development in these areas: Cloud computing such as AWS, Azure, and/or GCP. Infrastructure as code tools such as terraform or crossplane. Developing applications in languages such as python, ruby, or go. Deploying and supporting applications in Kubernetes at scale. 🎁 Benefits Company-subsidized medical, dental, & vision plans 401(k) plan with company match Annual bonus Flexible PTO to encourage a healthy work/life balance (2 weeks STRONGLY encouraged!) Generous paid leave programs, including 16-week paid parental leave and disability benefits Workplace flexibility and modern work schedules focused on getting the job done, not hours clocked