SRE III
Entain plcPuneFull TimeEngineering
Remotely
datadogkubernetesansibleterraformgithub actionsprometheusgrafanajenkins
Job Description
📋 Description
- Execute site reliability activities for scalable, fault-tolerant cloud services.
- Collaborate with cross-functional teams to improve performance, reliability, and availability.
- Develop and troubleshoot large-scale distributed systems in on-prem and cloud environments.
- Deliver infrastructure as code to enhance availability, scalability, latency, and efficiency.
- Monitor support processing for early detection of issues and share SRE trends.
- Automate tooling and CI/CD pipelines to promote best practices.
🎯 Requirements
- Infrastructure as Code (IaC) – Terraform, Ansible, or Kubernetes.
- Monitoring and Observability – logs, metrics, traces.
- Security and Compliance – access control, encryption, audit logs.
- Incident Management – quick production issue handling.
- Performance Optimization – speed, latency, resource usage.
- CI/CD – automated deployments with safe rollbacks.
🎁 Benefits
- Competitive compensation and pension.
- 24 days annual leave + wellbeing and development days.
- Life assurance and income protection.
- Private healthcare and wellbeing support.
- Communication allowance and crèche expense support.
- Respect for equal opportunities and inclusive hiring.