Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
pythonsreobservabilitysite reliabilityprometheusgrafanaalertmanageropenmetrics
Job Description
📋 Description
- Define SLIs/SLOs and error budgets for ~70 components.
- Build telemetry for customer-hosted agents and cloud services.
- Extend instrumentation in Python, Go, and Rust.
- Consolidate dashboards and observability tooling.
- Implement multi-window burn-rate alerting and runbooks.
- Establish ownership, escalation, and on-call practices.
🎯 Requirements
- Extensive production engineering or SRE with SLO framework experience.
- Strong Python; read/modify Go or Rust for instrumentation.
- Hands-on time-series telemetry (Prometheus/OpenMetrics, Grafana, Alertmanager).
- Experience debugging distributed systems beyond Kubernetes.
- Production config mgmt and CI/CD (Ansible, GitLab CI, Jenkins).
- Telemetry for push-based, non-scraped systems; privacy concerns.
🎁 Benefits
- Fully remote with flexible hours.
- Paid vacation, holidays, and sick leave.
- Private medical insurance support, space reimbursement.
- Gym/sports stipend and professional development.
- Remote-first, async org with global time zones.
- Opportunity to define SRE function and culture.
Back to all jobs