Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustprometheusclickhousegrafanaalertmanageropenmetrics
Job Description
📋 Description
- Define SLIs/SLOs and ownership across ~70 components.
- Build telemetry for customer-hosted agents and cloud services.
- Instrument Python, Go, and Rust components.
- Consolidate dashboards and enable reliable observability tooling.
- Establish incident response and blameless postmortems.
🎯 Requirements
- Extensive production SRE/engineering experience with SLO framework.
- Strong Python; ability to modify Go or Rust for instrumentation.
- Hands-on telemetry at scale: Prometheus/OpenMetrics, Grafana, Alertmanager, ClickHouse or similar.
- Experience debugging distributed systems beyond Kubernetes-centric view.
- Practical config mgmt and CI/CD tooling (Ansible, GitLab CI, Jenkins, etc).
- Telemetry for unscannable systems: push-based collection, privacy considerations.
🎁 Benefits
- Fully remote work with flexible hours worldwide.
- 25 days vacation, 10 holidays, unlimited sick leave.
- Private medical insurance contribution; coworking space reimbursement.
- Gym/sports reimbursement; professional development opportunities.
- Opportunity to define SRE function and reliability culture.
Back to all jobs