Site Reliability Engineer - Telemetry
Kraken Digital Asset ExchangeRemotely
kubernetesterraformprometheusgrafanaopentelemetryterragruntpyroscopevictoriametrics
Job Description
📋 Description
- Operate and improve the shared telemetry platform for metrics, logs, traces, alerting, dashboards
- Maintain metrics collection, storage, querying, dashboards, and alerting with Prometheus stack
- Manage log pipelines using Vector, Splunk, and Loki; ensure reliability and throughput.
- Operate tracing and profiling with Grafana Alloy, Tempo, OpenTelemetry, Pyroscope.
- Deploy telemetry services with Terraform, Terragrunt, and multi-env orchestration.
- Troubleshoot missing data, slow queries, broken alerts, capacity issues.
🎯 Requirements
- 3+ years in SRE/Platform/Observability or similar production role.
- Experience with large-scale telemetry data (metrics/logs/traces).
- Prometheus or Prometheus-compatible stack experience.
- Distributed systems troubleshooting; data flow and capacity issues.
- Infra as Code experience, esp. Terraform, and CI/CD.
- Experience with container workloads (Nomad, Kubernetes or similar).
🎁 Benefits
- Culture-driven, globally minded tech environment.
- On-call and incident learnings shape platform improvements.
- Commitment to diversity and equal opportunity hiring.