Site Reliability Engineer II
BackblazeRemotely
linuxansibleterraformmonitoringprometheusgrafanaelkcatchpoint
Job Description
📋 Description
- Ensure stability and reliability of Backblaze services and infrastructure.
- Build automation to reduce manual toil and improve efficiency.
- Support incident response and on-call rotations to maintain system performance.
- Collaborate with engineering, product, and operations to embed reliability.
🎯 Requirements
- Solid Linux administration and troubleshooting.
- Experience with monitoring/alerting and incident response; SRE fundamentals (SLOs, SLIs, error
- Proficiency in scripting (Python, Bash, or Go).
- Familiarity with containers (Kubernetes, Docker) and microservices.
- Experience with CI/CD, IaC (Terraform, Ansible, Jenkins).
- Good problem-solving skills and willingness to learn new tech.
🎁 Benefits
- Work with cloud technologies, SaaS/distributed systems.
- Exposure to ITIL/OSS practices and SLO/SLA concepts.
- Collaborative culture with focus on reliability and learning.