Research Engineer, LangSmith Engine
LangChainSan Francisco, New YorkFull TimeEngineering
Remotely
ai agentsllmsbenchmarksevaluationspost trainingexperiments
Job Description
📋 Description
- Build and maintain benchmarks and evaluations that measure quality and efficiency.
- Design and run experiments to improve agent performance across models, prompting, context, tools
- Explore post-training and fine-tuning techniques to improve capabilities, quality or cost.
- Turn successful experiments into production improvements, measure impact, prevent regressions.
- Help define ML roadmap and mentor engineers with strong technical leadership.
🎯 Requirements
- 4+ years of experience in ML/AI research or related field.
- Master’s or Ph.D. in a relevant field.
- Hands-on experience with LLMs and AI agents, analyzing behavior and improving performance.
- Strong experience designing benchmarks, evaluations, and experiments for AI/ML systems.
- Strong software engineering skills, taking ideas from prototype to production impact.
- Maximum agency and strong research judgment; ability to handle ambiguity and communicate findings.
🎁 Benefits
- Medical, dental, and vision coverage.
- Flexible vacation policy.
- 401(k) plan.
- Life insurance.