Job Description
📋 Description Operate and maintain all Develocity instances and supporting services. Participate in follow-the-sun on-call rotation; own incident response and troubleshooting. Drive automation across deployment, upgrades, monitoring, self-healing, and recovery. Build and maintain observability for all managed services (logging, metrics, tracing, and alerting). Work with engineering teams to build reliability into features from the start. Run incident retrospectives, and own disaster recovery, backups, and business continuity. 🎯 Requirements 5+ years in SRE/DevOps operating production services at scale Strong Kubernetes experience in production environments Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2) Proficiency with observability tools (Prometheus, Grafana) and IaC (Terraform) Track record of incident management and response; 24/7 on-call rotations Knowledge of SRE best practices (SLAs, SLOs) 🎁 Benefits A ground-floor role in a new SRE team Real ownership of production systems used by engineers at companies you've heard of Direct interaction with customers during incidents A culture that rewards automation over heroics Work from home in a remote-first environment Competitive salaries and equity grants