Senior Observability & Telemetry Engineer
JobgetherRemotely
gopythonrustkubernetesprometheusclickhousegrafanaopentelemetry
Job Description
📋 Description
- Design, operate scalable telemetry pipelines for metrics, logs, and traces across distributed GPU
- Architect telemetry storage for large-scale time-series and events; define instrumentation
- Build observability across compute, storage, networking, GPU clusters, training/inference workloads.
- Develop dashboards and tools for actionable workload health insights.
- Maintain network/infrastructure telemetry using Python or Go with gNMI, SNMP, and streaming
- Develop alerting, anomaly detection, reliability metrics, SLIs, and SLOs.
🎯 Requirements
- Proven experience operating observability systems at production scale with metrics, logging
- Strong programming in Go, Python, or Rust; experience building telemetry collectors, exporters
- Hands-on with Prometheus, OpenTelemetry, Grafana, distributed logging, and large-scale telemetry
- Experience with large-scale GPU cloud/HPC/AI infra and monitoring distributed workloads.
- Knowledge of GPU telemetry tech (NVIDIA DCGM, NVML, NVLink, NVSwitch) and AI workload metrics
- Understanding of networking/infrastructure telemetry (gNMI, SNMP, streaming telemetry, RDMA/RoCE).
🎁 Benefits
- Attractive compensation reflecting expertise and market conditions.
- Permanent, full-time role with an EMEA-based remote model.
- Flexible/hybrid-friendly environment for international collaboration.
- Work on large-scale GPU/AI/cloud/infrastructure challenges.
- Exposure to advanced observability and distributed-systems tech.
- Opportunity to influence observability standards and lead initiatives.
Back to all jobs