Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustansibleprometheusclickhousegrafanaopenmetrics
Job Description
📋 Description
- Define SLIs/SLOs for ~70 components with squads and engineers
- Build a reliability taxonomy across availability, latency, telemetry
- Ensure alerts have owners and actionable runbooks
- Develop telemetry for customer-hosted agents and cloud services
- Extend instrumentation in Python, Go, Rust
- Consolidate dashboards into a reliable observability platform
🎯 Requirements
- Production engineering or SRE with SLO framework experience
- Strong Python; read/modify Go or Rust for instrumentation
- Hands-on with time-series telemetry: Prometheus/OpenMetrics, Grafana
- Experience debugging distributed systems beyond Kubernetes
- CI/CD tooling: Ansible, GitLab CI, Jenkins
- Telemetry for non-scrapable systems; privacy considerations
🎁 Benefits
- Fully remote work with global hours
- 24 vacation days; 10 holidays; unlimited sick leave
- Private medical insurance contribution; co-working space reimbursement
- Gym/sports reimbursement; professional development
- Opportunity to define SRE function and culture
Back to all jobs