AI Research Resident
PolymathNorth AmericaContractEngineering
Remotely
systems engineeringreinforcement learningbenchmarkingautonomous agentsfrontier modelssimulation environmentsproduction quality codelong horizon reasoning
Job Description
📋 Description Collaborate on frontier benchmarks and environments for long-horizon AI agents. Identify failure modes in frontier models. Develop rigorous benchmarks for complex, long-horizon tasks and tool use. Train autonomous agents to reason, plan, and act over extended horizons. Offer full-time or part-time engagements. 🎯 Requirements Pursuing MS or PhD in CS or related field. Experience with reinforcement learning, benchmarking frontier models, or model post-training. Experience with systems engineering and production-quality code. Strong track record of publications. High agency, fast movers, and open-ended research problems.