Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
pythonsresite reliabilityprometheusclickhousegrafanasloopenmetrics
Job Description
📋 Description
- Define SLIs/SLOs for ~70 components with ownership and error budgets.
- Build telemetry and reliability platform for cloud and agent-based services.
- Instrument Python, Go, and Rust components for reliability.
- Consolidate dashboards; retire non-value tooling.
- Establish alerting with multi-window burn-rate; ensure actionable runbooks.
- Lead blameless postmortems and on-call improvements across time zones.
🎯 Requirements
- Production engineering or SRE with SLO framework experience.
- Strong Python; reading/modifying Go or Rust for instrumentation.
- Hands-on time-series telemetry: Prometheus/OpenMetrics, Grafana, Alertmanager, ClickHouse or
- Experience debugging distributed systems beyond Kubernetes.
- CI/CD tooling experience: Ansible, GitLab CI, Jenkins, etc.
- Telemetry for non-scrapable systems; privacy considerations on customer infra.
🎁 Benefits
- Fully remote work with global flexibility.
- 25 days vacation, paid holidays, unlimited sick leave.
- Private medical insurance contributions; coworking space reimbursement.
- Gym/sports reimbursement; professional development opportunities.
- Opportunity to define SRE function and operational culture.
Back to all jobs