Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustansibleprometheusclickhousegrafanaopenmetrics
Job Description
📋 Description
- Define SLIs/SLOs and error budgets for ~70 components.
- Build telemetry for customer-hosted agents and cloud services.
- Extend instrumentation in Python, Go, Rust.
- Consolidate dashboards and retire ineffective tooling.
- Implement SRE practices: alerting, escalation, on-call ownership.
- Lead reliability culture and measurable outcomes across teams.
🎯 Requirements
- Substantial production engineering or SRE experience with SLO framework design.
- Strong Python; ability to read/modify Go or Rust for instrumentation.
- Experience with time-series data at scale: Prometheus/OpenMetrics, Grafana, Alertmanager
- Experience debugging distributed systems beyond Kubernetes-centric view.
- Hands-on with CI/CD tooling (Ansible, GitLab CI, Jenkins).
- Telemetry for push-based collection, privacy, and data quality on customer-managed infra.
🎁 Benefits
- Fully remote work with flexible hours (worldwide).
- Competitive compensation and private medical insurance support.
- Co-working space and gym reimbursements.
- Professional development opportunities and mentorship.
- Remote-first, asynchronous environment across time zones.
- Opportunity to shape SRE function and operational culture.
Back to all jobs