Observability Specialist
JobgetherRemotely
datadogpostgresqlprometheusgrafanaopentelemetrysentryhoneycombpyroscope
Job Description
📋 Description
- Own, operate, and evolve observability platform across apps, APIs, databases, Kubernetes, cloud
- Improve production visibility with metrics, logs, traces, dashboards, alerts, and SLIs.
- Collaborate with engineering/platform teams to define SLIs and improve alert quality, reducing
- Build internal tools, libraries, automation, and self-service workflows for instrumenting services
- Identify reliability, latency, capacity, performance, and infra cost issues before user-facing
- Establish observability standards, naming conventions, docs, and training across engineering teams.
🎯 Requirements
- 5+ years in observability, SRE, platform eng, infra eng, backend eng, or production systems
- Strong understanding of metrics, logging, distributed tracing, profiling, alerting, dashboards
- Experience building internal tools, automation, libraries, or platforms used by other engineers.
- Proficiency in at least one backend language, preferably Python.
- Hands-on experience with observability platforms such as Prometheus, Grafana, Datadog
- Experience with OpenTelemetry instrumentation and collector configuration at scale.
🎁 Benefits
- Remote work opportunity with flexibility and autonomy.
- Opportunity to work on large-scale, mission-driven financial technology infrastructure.
- High ownership across projects from problem definition to production monitoring.
- Professional growth through exposure to observability, reliability, performance, and platform
- Opportunities to mentor and collaborate with experienced engineers across distributed teams.
- Inclusive and diverse international engineering environment.