AI Observability Engineer
JobgetherJob Description
📋 Description As an AI Observability Engineer, you will own the tools, processes, and systems that provide Design, implement, and operate observability solutions for AI workloads, including LLM and agent Configure and maintain AI tracing systems to capture latency, token usage, cost, quality metrics Develop internal tooling and automation solutions using Python for instrumentation and data Build and maintain dashboards, metrics, and monitoring solutions using Grafana, Prometheus, and Instrument AI platforms and workloads to provide visibility into health, usage, performance, cost 🎯 Requirements The ideal candidate is an experienced observability, SRE, DevOps, or cloud engineering professional 5–8 years of experience in observability, Site Reliability Engineering, platform engineering Strong hands-on experience with Azure Monitor, Application Insights, Log Analytics, and Managed Experience working with Langfuse, Grafana, and Prometheus for AI and application monitoring. Solid knowledge of Terraform and CI/CD practices. Strong Python skills for automation, instrumentation, exporters, and internal tooling development. 🎁 Benefits Competitive compensation package. Career development and continuous learning opportunities. Flexible working environment with strong ownership and autonomy. Opportunity to work on impactful AI infrastructure and technology projects. Collaborative culture with highly skilled international teams. Chance to contribute to the evolution of next-generation AI platforms.