San FranciscoFull TimeEngineering
Remotely
sqlfailure analysisnosqlevaluationllmsapisbenchmarksrubrics
Job Description
📋 Description
- Benchmarking: Design, implement, and maintain benchmarks and metrics for tool use, agentic
- Evaluation systems: Build and operate LLM evaluation systems end-to-end runs, scoring, dashboards
- Failure analysis: Run systematic failure analysis on model outputs (e.g., wrong tool use, reasoning
- Rubrics and evaluators: Create and refine rubrics, automated evaluators, and scoring frameworks
- Data quality and usability: Quantify data usability, quality, and impact on key benchmarks; use
- Cross-team collaboration: Work with AI researchers, applied AI teams, and data producers to align
🎯 Requirements
- Strong applied research background, with focus on model evaluation, benchmarking, and/or failure
- Strong coding skills and hands-on experience with ML models and evaluation code.
- Solid grasp of data structures, algorithms, and backend systems.
- Comfort with APIs, SQL/NoSQL, and cloud platforms for running and storing eval results.
- Ability to reason about model behavior, experimental results, and data quality from evals and
- Excitement to work in person in San Francisco five days a week in a high-intensity, high-ownership
🎁 Benefits
- Bi-annual performance bonus structure
- Generous equity grant vested over 4 years
- Up to $15k Relocation bonus
- $10K housing bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Free Equinox membership
Back to all jobs