Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustprometheusclickhousegrafanasentryopenmetrics
Job Description
📋 Description
- Define SLIs, SLOs, and error budgets for ~70 components with squad leads and senior engineers.
- Develop a reliability taxonomy covering availability, latency, telemetry health, and
- Ensure indicators are independently measurable and not disabled by the triggering failure.
- Design telemetry for customer-hosted agents and cloud services; balance push collection, sampling
- Extend instrumentation across Python, Go, and Rust components.
- Consolidate dashboards and reporting into a reliable observability platform; retire non-value
🎯 Requirements
- Substantial production engineering or SRE experience; experience defining and implementing an SLO
- Strong Python skills; ability to read/modify Go or Rust for instrumentation.
- Hands-on telemetry experience at scale: Prometheus/OpenMetrics, Grafana, Alertmanager, ClickHouse
- Experience debugging distributed systems beyond Kubernetes-centric view.
- Practical experience with production CI/CD tooling (Ansible, GitLab CI, Jenkins, etc.).
- Understanding telemetry for systems that cannot be fully scraped, including push-based collection
🎁 Benefits
- Fully remote work with flexible hours, work from anywhere worldwide.
- 24 paid vacation days per year.
- 10 paid national holidays.
- Unlimited sick leave.
- Private medical insurance contribution.
- Co-working space reimbursement.
Back to all jobs