Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustprometheusclickhousegrafanaalertmanageropenmetrics
Job Description
📋 Description
- Define SLIs for ~70 components with ownership, measurement, SLOs, and error budgets.
- Develop a reliability taxonomy for availability, latency, telemetry, and security control efficacy.
- Ensure reliability indicators are independent and not bypassable by failures they detect.
- Design telemetry collection for customer-hosted agents and cloud services; balance push-based
- Extend instrumentation across Python, Go, and Rust components with product teams.
- Consolidate dashboards into a reliable observability platform; retire non-value tooling.
🎯 Requirements
- Substantial production engineering or SRE experience with SLO framework ownership.
- Strong Python skills; ability to modify Go or Rust for instrumentation.
- Experience with time-series/event telemetry at scale (Prometheus/OpenMetrics, Grafana
- Hands-on debugging of distributed systems beyond Kubernetes-centric ops.
- Experience with production CI/CD tooling (Ansible, GitLab CI, Jenkins).
- Understanding of telemetry for systems with push-based collection, sampling, and privacy
🎁 Benefits
- Fully remote work with flexible hours (worldwide).
- 24 vacation days/year; 10 national holidays; unlimited sick leave.
- Private medical insurance contribution; coworking space reimbursement; gym rebate.
- Professional development through challenging projects and mentoring.
- Opportunity to earn a patent-worthy idea reward.
- Remote-first, asynchronous environment across time zones.