Senior Observability & Telemetry Engineer - Radian Arc
JobgetherRemotely
gopythonrustprometheusgrafanaopentelemetrysnmpgnmi
Job Description
📋 Description
- Design, implement, and operate scalable telemetry pipelines for metrics, logs, and traces across
- Architect and maintain telemetry storage for large-scale time-series and event data.
- Build observability across compute, storage, networking, GPU clusters, and AI workloads.
- Develop dashboards and tools for actionable workload health insights.
- Build and maintain network/infrastructure telemetry using Python or Go.
- Develop alerting, anomaly detection, SLIs, and SLOs for operational workflows.
🎯 Requirements
- Proven experience operating observability systems at production scale.
- Strong programming in Go, Python, or Rust; telemetry tooling experience.
- Hands-on with Prometheus, OpenTelemetry, Grafana; large-scale telemetry databases (ClickHouse).
- Experience with GPU cloud/HPC/AI infra and distributed workloads.
- GPU telemetry tech: DCGM, NVML, NVLink, NVSwitch; metrics for latency/throughput.
- Networking infra telemetry with gNMI, SNMP, streaming telemetry.
🎁 Benefits
- Attractive compensation package based on expertise.
- Permanent, full-time role with an EMEA-based remote model.
- Flexible, hybrid-friendly environment for international collaboration.
- Work on GPU/AI/cloud/networking/edge infrastructure challenges.
- Exposure to advanced observability and reliability tech.
- Mentorship and leadership opportunities across teams.
Back to all jobs