San FranciscoFull TimeEngineering
Remotely
llmdata scienceaievaluationproduction systemslangsmithcalibrationadversarial testing
Job Description
📋 Description
- Build testing harnesses and evaluation infrastructure for our agentic products
- Own evals end to end—architecture and content—with clinical partners
- Ensure deployments go through testing before shipping to patients
- Create ground truth, judges, and metrics; calibrate to human labels
- Develop labeling/review tooling for clinicians and designers
- Iterate evaluation results into model/prompt refinement safely
🎯 Requirements
- Significant experience shipping production systems and tooling
- Strong understanding of AI evaluation (LLM, red-teaming, synthetic scenarios)
- Experience deploying evals and automated AI tooling in production
- Ability to design workflows/UI for non-engineers
- Statistical fluency for calibration and reliability
- Interest in healthcare and interdisciplinary collaboration
🎁 Benefits
- Equal Opportunity Employer
- Opportunity to impact AI health coaching