Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustkubernetesprometheusclickhousegrafanaopentelemetry
Job Description
📋 Description
- Define SLIs for ~70 components with ownership, SLOs, and error budgets.
- Build reliability taxonomy for availability, telemetry, and config convergence.
- Ensure reliability indicators are independently measurable.
- Design telemetry pipelines for agents and cloud services.
- Extend instrumentation across Python, Go, Rust code.
- Consolidate dashboards and retire non-value tooling.
🎯 Requirements
- Significant production engineering or SRE exp with SLO framework work.
- Strong Python; ability to read/modify Go or Rust for instrumentation.
- Experience with time-series telemetry: Prometheus/OpenMetrics, Grafana, Alertmanager, ClickHouse.
- Debugging distributed systems beyond Kubernetes.
- Experience with Ansible, GitLab CI, Jenkins, or similar.
- Telemetry for non-scrapable/masked systems; privacy considerations.
🎁 Benefits
- Fully remote work with flexible hours (worldwide).
- 30+ days vacation equivalent (24 days) per year.
- 10 paid national holidays.
- Unlimited sick leave.
- Private medical insurance contribution.
- Co-working space reimbursement.
Back to all jobs