Remotely
customer servicedata analysisstatisticsdataevaluationmlrlbenchmark
Job Description
📋 Description
- Propose and scope a new benchmark or evaluation technique in a domain APEX doesn’t yet cover well
- Design task specifications and grading rubrics with Mercor’s network of domain experts.
- Build and validate the benchmark: pilot tasks, calibrate scoring, and test for contamination.
- Run frontier models against your benchmark and analyze failures.
- Publish results as papers, datasets, leaderboards, or methodologies for internal adoption.
- Collaborate with Mercor’s research and engineering teams to integrate learnings into APEX.
🎯 Requirements
- Genuine interest in evaluation as a research discipline.
- Background in CS, ML, statistics, or adjacent fields; no ML publication requirement.
- A specific, well-scoped idea for a benchmark or eval technique.
- Comfortable in a startup environment with fast iteration.
- Commit at least 20 hours/week for the duration of the fellowship.
- Bonus: experience with agentic evaluation, RL environments, or domain expertise in
🎁 Benefits
- 3 month stipend of $40,000 or 6 month stipend of $80,000.
- Unlimited API credits, GPU compute budget, and paid expert/human-data time.
- Weekly mentorship with the APEX research team and access to broader research org.
- Access to frontier model APIs and enterprise evaluation problems.
- Optional desk in Mercor’s San Francisco office; introductions to researchers across frontier labs
- High-potential fellows may receive a full-time offer after the fellowship.