Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
sresite reliabilityci/cdprometheusclickhousegrafanaopentelemetryslo
Job Description
📋 Description
- Define SLIs for ~70 components with ownership and SLOs.
- Develop reliability taxonomy for availability, latency, telemetry.
- Ensure indicators are measurable and cannot be disabled.
- Design telemetry pipeline for agents and cloud services.
- Extend instrumentation in Python, Go, Rust.
- Consolidate dashboards into a unified observability platform.
🎯 Requirements
- Substantial prod engineering or SRE experience with an SLO framework.
- Strong Python skills; read/modify Go or Rust for instrumentation.
- Experience with Prometheus/OpenMetrics, Grafana, Alertmanager, ClickHouse or similar.
- Debug distributed systems on bare metal/long-lived hosts beyond Kubernetes.
- Production CI/CD tooling: Ansible, GitLab CI, Jenkins, or similar.
- Telemetry for systems not directly scraped: push-based collection, sampling, privacy.
🎁 Benefits
- Fully remote work with flexible hours worldwide.
- 24 paid vacation days per year.
- 10 paid national holidays.
- Unlimited sick leave.
- Private medical insurance reimbursement.
- Co-working space reimbursement.
Back to all jobs