Lead Site Reliability Engineer - Imunify Reliability Platform
JobgetherRemotely
gopythonrustprometheusclickhousegrafanaalertmanageropenmetrics
Job Description
📋 Description
- Define SLIs/SLOs and ownership for ~70 components.
- Build telemetry pipelines for customer-hosted agents and cloud services.
- Extend instrumentation across Python, Go, and Rust.
- Consolidate dashboards and alerting into a reliable observability platform.
- Establish escalation, on-call, and postmortem practices across time zones.
- Shape reliability culture for a major security product.
🎯 Requirements
- Strong production engineering or SRE experience with SLO framework definition.
- Proficient in Python; ability to read/modify Go or Rust for instrumentation.
- Hands-on with time-series/event telemetry at scale (Prometheus/OpenMetrics, Grafana, Alertmanager
- Experience debugging distributed systems beyond Kubernetes-centric ops.
- Practical experience with CI/CD tooling (Ansible, GitLab CI, Jenkins, etc.).
- Telemetry for systems with push-based collection, privacy considerations, and partial reporting.
🎁 Benefits
- Fully remote work with flexible hours.
- 24 vacation days, 10 holidays, unlimited sick leave.
- Private medical insurance compensation, co-working space reimbursement, gym/fitness support.
- Professional development and opportunities to patent ideas.
- Remote-first, asynchronous working across time zones.
- Opportunity to define an SRE function and reliability culture from the ground up.
Back to all jobs