Senior Site Reliability Engineer
GradleJob Description
📋 Description Operate and maintain all Develocity instances and supporting services. Participate in a follow-the-sun on-call rotation, owning incident response and troubleshooting across the stack. Drive automation across application deployment, upgrades, monitoring, self-healing, and recovery. Build and maintain observability for all managed services (logging, metrics, tracing, and alerting). Work with engineering teams to build reliability into features from the start. Run incident response and retrospectives, and ensure learning from outcomes. 🎯 Requirements 5+ years in SRE, DevOps, or equivalent role operating production services at scale. Strong Kubernetes experience in production environments. Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2). Proficiency with observability tools (Prometheus, Grafana) and Terraform. Track record of incident management and response. Knowledge of SRE best practices (SLAs, SLOs). 🎁 Benefits Ground-floor SRE role; shape processes and practices. Own production systems used by engineers at major companies. Direct customer interaction during incidents. Culture that values automation over heroics. In-person team offsites and annual meetings. Work from home in a remote-first environment.