Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustprometheusclickhousegrafanaopentelemetryalertmanager
Job Description
📋 Description
- Lead SRE for Imunify Reliability Platform across ~70 components
- Define SLIs, SLOs, error budgets, monitoring, alerting, and escalation
- Build telemetry pipeline for cloud services and agent-based deployments
- Collaborate with engineering teams to extend instrumentation (Python, Go, Rust)
- Establish blameless postmortems and ownership models
- Shape reliability culture with measurable outcomes
🎯 Requirements
- Strong production engineering or SRE background with SLO framework experience
- Proficient in Python; able to read/modify Go or Rust for instrumentation
- Hands-on with time-series telemetry at scale (Prometheus/OpenMetrics, Grafana, Alertmanager)
- Experience debugging distributed systems beyond Kubernetes-centric scope
- Practical with CI/CD tooling (Ansible, GitLab CI, Jenkins)
- Understanding of push-based telemetry, privacy considerations, and data quality
🎁 Benefits
- Fully remote work with flexible hours
- 24 paid vacation days + 10 paid holidays
- Unlimited sick leave
- Private medical insurance contribution + coworking space reimbursement
- Gym/sports reimbursement and development opportunities
- Remote-first, async environment with growth potential
Back to all jobs