San FranciscoFull TimeEngineering
Remotely
pythonrustobservabilityprometheusclickhousegrafanaopentelemetrytelemetry
Job Description
📋 Description
- Design and build scalable telemetry ingest and storage pipelines for metrics, logs, and traces
- Own and evolve core observability platforms, driving migrations and architectural improvements that
- Build instrumentation libraries, SDKs, and integrations that emit high-quality telemetry from
- Drive alerting and SLO infrastructure enabling reliable monitoring with minimal noise
- Collaborate with Inference, Product, and Infrastructure teams to meet diverse observability needs
🎯 Requirements
- Deep experience in at least one observability signal area (metrics, logging, tracing, or error
- Understanding of high-throughput data pipelines, columnar storage engines, and telemetry data at
- Experience with observability platforms such as Prometheus, Grafana, ClickHouse, OpenTelemetry, or
- Strong proficiency in Python, Rust, or Go
- Excellent communication and collaboration skills for improving operational visibility and incident
- Interest in applying AI/LLMs to operational workflows (e.g., automated root cause analysis, anomaly
🎁 Benefits
- Competitive compensation with meaningful equity
- 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO including Winter Break
- Paid parental leave
- Fertility and family-building stipend
- Company-facilitated 401(k)