Staff Site Reliability Engineer (AWS, Terraform, Distributed Systems)
JobgetherRemotely
gopythonawsdatadogterraformdistributed systemsprometheusgrafana
Job Description
📋 Description
- Architect, implement, and maintain highly available services with focus on resilient distributed
- Automate monitoring, deployment, operational processes, and incident-response to improve
- Lead complex troubleshooting, identify root causes, and prevent recurring incidents.
- Collaborate with global teams to build, deploy, maintain, and improve application services.
- Modernize traditional apps toward cloud-native architectures and influence key technical decisions.
- Design and manage scalable infrastructure using AWS, Terraform, and Ansible.
🎯 Requirements
- Bachelor’s degree with 5+ years or advanced degree with relevant experience; equivalent experience
- Strong SRE/DevOps experience with cloud infra and distributed systems.
- Hands-on with AWS, exposure to Azure or GCP; hybrid cloud a plus.
- Proven IaC experience with Terraform and/or Ansible.
- Experience deploying containerized microservices and high-performance distributed apps.
- Strong observability skills with Prometheus, Grafana, Datadog, ELK.
🎁 Benefits
- Fully remote within India with possible occasional office attendance with notice.
- Opportunity to work on large-scale distributed systems and globally significant tech platforms.
- Exposure to modern cloud, IaC, observability, automation, containerization, and ML tech.
- Strong leadership and continuous learning opportunities.
- Collaborative environment with global teams and stakeholders.
- Experience with modern SRE practices in fast-paced, 24x7 settings.
Back to all jobs