Senior Observability & Telemetry Engineer - Radian Arc
JobgetherRemotely
gopythonrustprometheusclickhousegrafanaopentelemetrygnmi
Job Description
📋 Description
- Design, implement, and operate scalable telemetry pipelines for metrics, logs, and traces across
- Architect telemetry storage for time-series and event data; define instrumentation and SLIs/SLOs.
- Build observability across compute, storage, networking, GPU, and AI workloads; identify
- Develop dashboards and monitoring tools for workload health and performance insights.
- Build network/infra telemetry using Python or Go; integrate gNMI, SNMP, and streaming telemetry.
- Develop alerting and anomaly detection; integrate signals into incident workflows.
🎯 Requirements
- Proven experience operating observability systems at production scale.
- Strong Go, Python, or Rust programming skills.
- Hands-on with Prometheus, OpenTelemetry, Grafana; large telemetry DBs like ClickHouse.
- Experience with GPU cloud/HPC/AI infra and distributed training/inference monitoring.
- Knowledge of NVIDIA DCGM, NVML, GPU telemetry and AI workload metrics.
- Strong networking/infra telemetry experience (gNMI, SNMP, streaming telemetry).
🎁 Benefits
- Attractive compensation; permanent, full-time; EMEA-based remote.
- Flexible/hybrid-friendly environment for international collaboration.
- Work on GPU/AI/cloud infra challenges; exposure to advanced observability tech.
- Opportunity to influence platform standards and mentor engineers.
- Career growth in a fast-growing international scale-up.
- Inclusive, diverse environment with fair development support.
Back to all jobs