Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustansibleprometheusclickhousegrafanaopenmetrics
Job Description
📋 Description
- Define SLIs/SLOs and ownership for ~70 components across cloud services and agents.
- Build telemetry and reliability platform; improve detection of silent degradations.
- Extend instrumentation to Python, Go, and Rust codebases.
- Create a scalable alerting system with clear runbooks and ownership.
- Collaborate with engineering leads to shape reliability culture and foundations.
🎯 Requirements
- Substantial production engineering or SRE experience with SLO framework implementation.
- Strong Python skills; ability to modify Go or Rust for instrumentation.
- Experience with time-series telemetry at scale (Prometheus/OpenMetrics, Grafana, Alertmanager).
- Hands-on with distributed systems on bare metal/long-lived hosts, not just Kubernetes.
- Experience with CI/CD tools (Ansible, GitLab CI, Jenkins).
- Knowledge of telemetry for non-scrapable systems, privacy, and partial reporting.
🎁 Benefits
- Fully remote work with flexible hours worldwide.
- 24 vacation days; 10 paid holidays; unlimited sick leave.
- Private medical insurance contribution, coworking space reimbursement.
- Gym/sports reimbursement; professional development opportunities.
- Opportunity to define an SRE function and reliability culture from the ground up.
Back to all jobs