Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherNorth AmericaFull TimeEngineering
Remotely
gopythonrustprometheusclickhousegrafanaalertmanageropenmetrics
Job Description
📋 Description
- Define and establish meaningful SLIs for ~70 components with ownership, measurement, SLOs, and
- Develop reliability taxonomy for service availability, latency, telemetry health, and configuration
- Ensure reliability indicators are independently measurable and cannot be disabled by failures they
- Design telemetry collection for customer-hosted agents and cloud services with privacy, data
- Extend instrumentation across Python, Go, and Rust components.
- Consolidate dashboards and reporting into a lean observability platform; retire low-value tooling.
🎯 Requirements
- Substantial production engineering or SRE experience with a defined SLO framework.
- Strong Python skills; ability to modify Go or Rust for instrumentation.
- Hands-on time-series/event telemetry experience at scale (Prometheus/OpenMetrics, Grafana
- Experience debugging distributed systems beyond Kubernetes-centric ops.
- Practical experience with production CI/CD tooling (Ansible, GitLab CI, Jenkins).
- Understanding of telemetry for systems with push-based collection, sampling, clock skew, and
🎁 Benefits
- Fully remote work with flexible hours worldwide.
- 24 paid vacation days per year.
- 10 paid national holidays.
- Unlimited sick leave.
- Contribution toward private medical insurance.
- Co-working space reimbursement.
Back to all jobs