AI Observability Engineer
JobgetherJob Description
📋 Description Own tools, processes, and systems for visibility into AI applications, platform services, and Design, implement, and operate observability solutions for AI workloads, including LLM and agent Configure and maintain AI tracing systems to capture latency, token usage, cost, quality metrics Develop internal tooling and automation solutions using Python for instrumentation and data Build and maintain dashboards, metrics, and monitoring solutions using Grafana, Prometheus, and Instrument AI platforms to provide visibility into health, usage, performance, cost, and SLIs/SLOs. 🎯 Requirements 5–8 years of experience in observability, SRE, platform engineering, DevOps, or cloud engineering Strong hands-on experience with Azure Monitor, Application Insights, Log Analytics, and Managed Experience with Langfuse, Grafana, and Prometheus for AI and application monitoring. Solid knowledge of Terraform and CI/CD practices. Strong Python skills for automation, instrumentation, exporters, and internal tooling development. Familiarity with machine learning workloads and AI-specific observability requirements. 🎁 Benefits Competitive compensation package. Career development and continuous learning opportunities. Flexible working environment with strong ownership and autonomy. Opportunity to work on impactful AI infrastructure and technology projects. Collaborative culture with highly skilled international teams. Chance to contribute to the evolution of next-generation AI platforms.