Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustansibleprometheusclickhousegrafanaopenmetrics
Job Description
📋 Description
- Define SLIs/SLOs and ownership across ~70 components with squads and senior engineers.
- Build reliability taxonomy for availability, latency, telemetry health, and configuration
- Design telemetry collection for customer-hosted agents and cloud services with privacy and data
- Extend instrumentation to Python, Go, and Rust; consolidate dashboards into a reliable
- Implement SLO-driven alerting with multi-window burn-rate and clear classifications.
- Establish owner definitions and runbooks for production alerts; ensure blameless postmortems and
🎯 Requirements
- Substantial production engineering or SRE experience with SLO framework definition/implementation.
- Strong Python skills; ability to modify Go or Rust for instrumentation.
- Hands-on time-series/event telemetry experience (Prometheus/OpenMetrics, Grafana, Alertmanager
- Experience debugging distributed systems beyond Kubernetes-centric ops.
- Practical CI/CD tooling experience (Ansible, GitLab CI, Jenkins).
- Telemetry for non-scrapable systems; privacy considerations on customer-managed infra.
🎁 Benefits
- Fully remote work with flexible hours; work from anywhere worldwide.
- 24 paid vacation days/year; 10 paid holidays; unlimited sick leave.
- Private medical insurance contribution; coworking space and gym reimbursements.
- Professional development opportunities through projects, mentoring, knowledge sharing.
- Opportunity to patent innovative ideas; remote-first, asynchronous environment.
- Chance to define SRE function, reliability standards, and culture from the ground up.
Back to all jobs