Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustkubernetesprometheusclickhousegrafanaopenmetrics
Job Description
📋 Description
- Define SLIs/SLOs and ownership for ~70 components
- Build telemetry and reliability platform for cloud and customer-hosted agents
- Instrument Python, Go, and Rust components for reliability
- Establish scalable alerting, runbooks, and on-call processes across time zones
- Collaborate with engineering leads to improve detection of security-control degradation
🎯 Requirements
- Substantial production engineering or SRE experience with SLO framework definition
- Strong Python skills; ability to read/modify Go or Rust for instrumentation
- Hands-on experience with time-series/event telemetry at scale (Prometheus/OpenMetrics, Grafana
- Experience debugging distributed systems beyond Kubernetes-centric ops
- Experience with CI/CD tooling (Ansible, GitLab CI, Jenkins)
- Understanding of telemetry for non-scraped systems, including push-based collection, sampling
🎁 Benefits
- Fully remote work with flexible hours
- 24 vacation days, 10 holidays, unlimited sick leave
- Private medical insurance contribution, co-working space reimbursement
- Gym and sports reimbursement, professional development opportunities
- Opportunity to define SRE function and reliability culture
Back to all jobs