Job Description
📋 Description Design and own evaluation harnesses for LLM and agentic outputs. Build and calibrate LLM-as-judge pipelines; validate against human labels and bias. Define and track response-quality metrics: fidelity, grounding, hallucination. Curate, version, and grow evaluation datasets as the product surfaces evolve. Benchmark models to allocate tasks and quantify quality vs compute. Build production forecasting, anomaly detection, and early-warning signals from telemetry. 🎯 Requirements Bachelor's or Master's in Data Science, CS, Statistics, Math, or related field or equivalent experience. 3+ years in data science, ML, or applied quantitative analysis. Strong applied statistics with experimental design and significance testing for noisy outputs. Experience building, deploying, and monitoring predictive or time-series models in production. Demonstrated work evaluating, analyzing, or improving LLM/NLP systems: eval design, retrieval eval, or agent analysis. Proficiency in Python and SQL; fluency with foundation models and LLM tooling. 🎁 Benefits Comprehensive medical, dental and vision plans. 401(k) plan with employer match. Flexible PTO and Volunteer Time Off (VTO). 5-year Service Milestone Sabbatical. Paid parental leave. Generous employee referral bonus program.