Senior Observability & Telemetry Engineer - Radian Arc
JobgetherRemotely
gopythonrustprometheusclickhousegrafanaopentelemetrygnmi
Job Description
📋 Description
- Design, implement, and operate telemetry pipelines for metrics, logs, and traces across distributed
- Architect telemetry storage for large-scale time-series and event data.
- Build observability across compute, storage, networking, GPU clusters, and AI workloads.
- Develop dashboards and monitoring tools for workload health and performance.
- Maintain network and infra telemetry using Python or Go.
- Develop alerting, anomaly detection, SLIs, and SLOs.
🎯 Requirements
- Experience operating observability systems at production scale with metrics, logging, tracing, and
- Strong Go, Python, or Rust skills for telemetry tooling.
- Hands-on with Prometheus, OpenTelemetry, Grafana, and large-scale telemetry DBs like ClickHouse.
- Experience with GPU cloud/HPC/AI infra and monitoring distributed training.
- Knowledge of GPU telemetry tech (NVIDIA DCGM, NVML, etc.) and AI metrics (inference latency
- Strong networking/infra telemetry know-how (gNMI, SNMP, streaming telemetry, RDMA/RoCE).
🎁 Benefits
- Attractive compensation package.
- EMEA-based remote; permanent, full-time.
- Hybrid-friendly environment for international collaboration.
- Work on large-scale GPU, AI, cloud, networking challenges.
- Exposure to advanced observability and distributed-systems tech.
- Mentorship and leadership opportunities.
Back to all jobs