Observability Platform Engineer — Neocloud
MirantisRemotely
gopythonkubernetesprometheusgrafanaopentelemetrylokijaeger
Job Description
📋 Description
- Design, build, and operate observability components for large-scale infrastructure
- Develop telemetry pipelines with high cardinality and volume while balancing cost and performance
- Define SLO/SLI frameworks and alerting to surface real signal
- Collaborate with operations teams to deliver what they need during incidents
- Improve detection speed and MTTR/MTTD across the platform
🎯 Requirements
- Proven experience designing and building observability platforms at scale
- Hands-on with metrics, logging, and tracing tooling (Prometheus, Grafana, OpenTelemetry, Loki
- Experience with high-volume telemetry pipelines and tradeoffs (cardinality, retention, cost
- Strong software engineering in Go, Python, or Rust
- Experience with Kubernetes and cloud-native infrastructure
- Understanding of SLO/SLI and noise-reduction in alerting
🎁 Benefits
- Competitive compensation and benefits
- Professional development, conferences, and hackathons
- Open-source/open standards focus with opportunities to contribute