Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustansibleprometheusclickhousegrafanaopenmetrics
Job Description
📋 Description
- Lead SRE for Imunify Reliability Platform in a remote setting
- Define SLIs/SLOs and error budgets across ~70 components
- Build a reliability taxonomy for availability, latency, telemetry, and security
- Architect telemetry collection for cloud services and agent-based workloads
- Instrument Python, Go, and Rust components for reliability gains
- Unify dashboards and reporting into a streamlined observability platform
🎯 Requirements
- Production SRE or engineering experience with SLO framework creation
- Strong Python; reading/modifying Go or Rust for instrumentation
- Experience with time-series telemetry: Prometheus/OpenMetrics, Grafana, Alertmanager
- Hands-on with scalable telemetry in colum nar/high-cardinality stores (e.g., ClickHouse)
- Debugging distributed systems beyond Kubernetes-centric scope
- CM/CI tooling: Ansible, GitLab CI, Jenkins
🎁 Benefits
- Fully remote, flexible hours, work from anywhere
- 24 vacation days/year
- 10 paid national holidays
- Unlimited sick leave
- Private medical insurance contribution
- Co-working space reimbursement
Back to all jobs