Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustsreprometheusclickhousegrafanaopenmetrics
Job Description
📋 Description
- Define SLIs/SLOs and ownership for ~70 components.
- Build reliability taxonomy across availability, latency, telemetry health.
- Ensure measurable, un-disablable reliability indicators.
- Design telemetry pipelines for agent and cloud services, balancing privacy & data quality.
- Extend instrumentation across Python, Go, Rust components.
- Consolidate dashboards and reporting into a focused observability platform.
🎯 Requirements
- Substantial production engineering or SRE experience with SLO framework ownership.
- Strong Python; ability to read/modify Go or Rust for instrumentation.
- Experience with time-series/event telemetry at scale: Prometheus/OpenMetrics, Grafana
- Hands-on debugging of distributed systems beyond Kubernetes-centrics.
- Production config management and CI/CD tooling (Ansible, GitLab CI, Jenkins, etc.).
- Telemetry for systems with push-based collection, sampling, clock skew, privacy concerns.
🎁 Benefits
- Fully remote with global flexibility.
- 24 vacation days, 10 holidays, unlimited sick leave.
- Private medical insurance contribution, coworking space reimbursement.
- Gym/sports reimbursement and professional development opportunities.
- Opportunity to patent innovative ideas; remote-first, asynchronous culture.
- Chance to shape SRE function and reliability culture from the ground up.
Back to all jobs