Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustprometheusclickhousegrafanaopentelemetryopenmetrics
Job Description
📋 Description
- Define SLIs/SLOs and error budgets for ~70 components.
- Build telemetry for customer-hosted agents and cloud services.
- Extend instrumentation across Python, Go, and Rust.
- Consolidate dashboards into a reliable observability platform.
- Implement SLO-driven alerting with multi-window burn-rate.
- Ensure production alerts have owners and runbooks.
🎯 Requirements
- Substantial production Eng/SRE experience with SLO framework.
- Strong Python; read/modify Go or Rust for instrumentation.
- Hands-on with time-series telemetry: Prometheus/OpenMetrics, Grafana.
- Experience debugging distributed systems, not just Kubernetes.
- Config management/CI/CD: Ansible, GitLab CI, Jenkins.
- Telemetry for systems with push-based reporting and privacy concerns.
🎁 Benefits
- Fully remote work with global hours.
- 24 vacation days per year.
- 10 paid national holidays.
- Unlimited sick leave.
- Private medical insurance support.
- Co-working space reimbursement.
Back to all jobs