Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustprometheusclickhousegrafanaopentelemetryalertmanager
Job Description
📋 Description
- Define SLIs for ~70 components with ownership, SLOs, and error budgets.
- Build telemetry pipeline for customer-hosted agents and cloud services.
- Extend instrumentation across Python, Go, and Rust components.
- Consolidate dashboards and reporting into a reliable observability platform.
- Implement SRE-focused alerting with clear escalation and runbooks.
- Establish blameless postmortems and cross-team reliability practices.
🎯 Requirements
- Production engineering or SRE experience with SLO framework design.
- Strong Python; ability to read/modify Go or Rust.
- Hands-on time-series telemetry: Prometheus/OpenMetrics, Grafana, Alertmanager, ClickHouse or
- Experience debugging distributed systems on bare metal/long-lived hosts (not just Kubernetes).
- CI/CD tooling experience (Ansible, GitLab CI, Jenkins).
- Telemetry for non-scrapable systems; privacy and partial reporting considerations.
🎁 Benefits
- Fully remote work with flexible hours (worldwide).
- 24 vacation days, 10 holidays, unlimited sick leave.
- Private medical insurance contribution, co-working and gym reimbursements.
- Professional development through challenging projects and mentoring.
- Option to define SRE function and operational culture from the ground up.
Back to all jobs