Job Description
📋 Description Own production services from design to day‑to‑day operation Mentor SRE I and II; lead reliability initiatives across the org Apply AI/LLM tools to speed incident response and automate runbooks Build scalable, highly available systems; ensure fast, resilient services Partner with product, engineering, and security teams on critical systems 🎯 Requirements 6+ years in Site Reliability Engineering or Platform Engineering Deep Kubernetes and Docker production experience Advanced Linux skills (kernel, TCP/IP, DNS, load balancers) Python automation with Bash scripting; experience with Ansible and IaC tools (Terraform or Pulumi) Hands-on with monitoring: Icinga, Prometheus, Grafana Familiarity with AI/ML concepts and integrating AI-ops tools 🎁 Benefits Competitive total rewards; benefits vary by location and role Be part of a diverse, inclusive culture with employee resource groups Remote-friendly with occasional office visits for events/meetings