Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustprometheusclickhousegrafanaalertmanageropenmetrics
Job Description
📋 Description
- Define SLIs/SLOs and error budgets for ~70 components
- Build telemetry for cloud and agent-hosted services
- Extend instrumentation across Python, Go, Rust
- Consolidate dashboards and reporting into a reliable platform
- Implement SLO-driven, symptom-based alerting
🎯 Requirements
- Production engineering or SRE experience with SLO framework work
- Strong Python; ability to modify Go or Rust for instrumentation
- Experience with time-series telemetry: Prometheus/OpenMetrics, Grafana, Alertmanager, ClickHouse or
- Debugging distributed systems beyond Kubernetes
- CI/CD tooling: Ansible, GitLab CI, Jenkins
- Telemetry for non-scraped systems, privacy considerations
🎁 Benefits
- Fully remote work, work from anywhere
- 24 vacation days, 10 holidays, unlimited sick leave
- Private medical insurance contribution and coworking space reimbursement
- Gym/sports reimbursement and professional development
- Remote-first, asynchronous, multi-time-zone environment
Back to all jobs