Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustprometheusclickhousegrafanaalertmanageropenmetrics
Job Description
📋 Description
- Define SLIs, SLOs, and ownership for ~70 components with squads and senior engineers.
- Develop a reliability taxonomy covering availability, latency, telemetry health, and configuration
- Ensure indicators are independently measurable and cannot be defeated by the failure they detect.
- Design telemetry collection for customer-hosted agents and cloud services with privacy and data
- Extend instrumentation across Python, Go, and Rust components.
- Consolidate dashboards and reporting into a reliable observability platform.
🎯 Requirements
- Substantial production engineering or SRE experience with SLO framework design and implementation.
- Strong Python skills; able to modify Go or Rust code for instrumentation.
- Hands-on experience with Prometheus/OpenMetrics, Grafana, Alertmanager, and ClickHouse or similar
- Experience debugging distributed systems beyond Kubernetes-centric ops.
- Practical experience with Ansible, GitLab CI, Jenkins, or similar CI/CD tooling.
- Strong telemetry understanding for non-scrapable systems, including privacy considerations.
🎁 Benefits
- Fully remote work with flexible hours (Worldwide).
- 24 vacation days/year + 10 paid holidays.
- Unlimited sick leave.
- Private medical insurance compensation and co-working space reimbursement.
- Gym/sports reimbursement and professional development opportunities.
- Remote-first, asynchronous work across time zones; opportunity to shape SRE function.
Back to all jobs