Job Description
📋 Description Role: Site Reliability Engineer in Core Platform Engineering. Build automated reliability tooling, IaC, and observability to improve uptime and performance. Develop monitoring, logging, and alerting frameworks (Prometheus, Grafana, OpenTelemetry). Implement automated architectural reviews and guardrails for scalable machine-generated apps. Partner with engineering to design scalable, fault-tolerant systems with defined SLIs/SLOs. Automate operational tasks and create self-healing and auto-remediation mechanisms. 🎯 Requirements Bachelor’s degree in Computer Science, Software Engineering, or related field. 3-5 years in Site Reliability Engineering, DevOps, or Infrastructure Eng. 2+ years in cloud environments (AWS, GCP, or Azure). 1+ years in at least one modern programming language (e.g., Java, Go, Python, Ruby, JavaScript). 🎁 Benefits Comprehensive medical, dental, vision, HSA/FSAs, life and AD&D insurance, 401(k) with match. Parental leave, unlimited PTO subject to policy, holidays, disability insurance, and more. Learning and development benefits, employee assistance, pet insurance, travel assistance, discounts.