Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustprometheusclickhousegrafanaalertmanageropenmetrics
Job Description
📋 Description
- Define SLIs, SLOs for ~70 components with ownership and error budgets.
- Develop reliability taxonomy for availability, latency, telemetry health.
- Ensure indicators are measurable and not spoofable by failures.
- Build telemetry pipeline for customer-hosted agents and cloud services.
- Extend instrumentation across Python, Go, Rust components.
- Consolidate dashboards and reporting into a lean observability platform.
🎯 Requirements
- Production engineering or SRE experience with SLO framework setup.
- Strong Python; able to read/modify Go or Rust code for instrumentation.
- Hands-on time-series/event telemetry at scale: Prometheus/OpenMetrics, Grafana, Alertmanager
- Experience debugging distributed systems beyond Kubernetes-centric ops.
- Practical config mgmt and CI/CD tooling (Ansible, GitLab CI, Jenkins).
- Telemetry for systems with push-based collection, privacy, and partial reporting.
🎁 Benefits
- Fully remote work with global flexibility.
- Professional development through challenging projects and mentoring.
- Remote-first, asynchronous environment across time zones.
- Opportunity to define an SRE function and reliability culture.
Back to all jobs