AI Observability Engineer
JobgetherJob Description
📋 Description Own tools, processes, and systems providing visibility into AI applications, platform services, and Design, implement, and operate observability solutions for AI workloads, including LLM and agent Configure and maintain AI tracing systems to capture latency, token usage, cost, quality metrics Develop internal tooling and automation using Python for instrumentation and data collection. Build and maintain dashboards, metrics, and monitoring solutions using Grafana, Prometheus, and Instrument AI platforms to provide visibility into health, usage, performance, cost, and SLOs 🎯 Requirements 5–8 years in observability, SRE, DevOps, or cloud engineering roles. Strong hands-on experience with Azure Monitor, Application Insights, Log Analytics, and Managed Experience with Langfuse, Grafana, and Prometheus for AI/application monitoring. Solid Terraform and CI/CD practices; strong Python skills for automation and tooling development. Familiarity with ML workloads and AI observability requirements; knowledge of logs, metrics Ability to work effectively in distributed, fast-paced, collaborative environments; intermediate or 🎁 Benefits Competitive compensation package. Career development and continuous learning opportunities. Flexible working environment with ownership and autonomy. Opportunity to work on impactful AI infrastructure projects with international teams. Collaborative culture in a highly skilled environment.