Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustansibleprometheusclickhousegrafanaopenmetrics
Job Description
📋 Description
- Define SLIs/SLOs for ~70 components with ownership, measurement, and error budgets.
- Build a reliability taxonomy covering availability, latency, telemetry health, and security
- Design telemetry pipelines for customer-hosted agents and cloud services.
- Extend instrumentation across Python, Go, and Rust components.
- Consolidate dashboards and establish a reliable observability platform.
- Implement SLO-driven, symptom-based alerting with clear runbooks.
🎯 Requirements
- Production engineering or SRE experience with SLO framework definition.
- Strong Python; ability to read/modify Go or Rust for instrumentation.
- Hands-on time-series/event telemetry exp. with Prometheus/OpenMetrics, Grafana, Alertmanager
- Experience debugging distributed systems beyond Kubernetes-centric ops.
- Practical config mgmt/CI-CD tooling (Ansible, GitLab CI, Jenkins).
- Understand telemetry for systems with push-based collection, privacy concerns on customer infra.
🎁 Benefits
- Fully remote with worldwide flexibility.
- 24 vacation days; 10 holidays; unlimited sick leave.
- Private medical insurance contribution; co-working space and gym reimbursements.
- Professional development and opportunities to patent ideas.
- Remote-first, asynchronous work across time zones.
- Opportunity to define SRE function and operational culture from the ground up.
Back to all jobs