Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
sresite reliabilityebpfprometheusclickhousegrafanaopentelemetryslo
Job Description
📋 Description
- Define SLIs/SLOs for ~70 product components with ownership, measurement, and error budgets.
- Build a reliability platform and telemetry pipeline for cloud services and customer-hosted agents.
- Extend instrumentation across Python, Go, and Rust; consolidate dashboards and observability
- Implement SLO-driven alerting with multi-window burn-rate and clear classifications.
- Establish escalation models, on-call practices, and blameless postmortems across time zones.
- Lead discussions to shape reliability culture and measurable outcomes.
🎯 Requirements
- Substantial production engineering or SRE experience with an SLO framework.
- Strong Python skills; ability to read/modify Go or Rust for instrumentation.
- Hands-on experience with Prometheus/OpenMetrics, Grafana, Alertmanager, ClickHouse (or similar).
- Experience debugging distributed systems beyond Kubernetes-centric operations.
- Practical experience with CI/CD tooling (Ansible, GitLab CI, Jenkins).
- Understanding of push-based telemetry, privacy, and data quality for customer-managed infra.
🎁 Benefits
- Fully remote work with flexible hours, globally.
- 28 days total vacation, 10 holidays, unlimited sick leave.
- Private medical insurance contribution, coworking space, gym reimbursements.
- Professional development and opportunities to shape an SRE function.
Back to all jobs