Job Description
📋 Description Design evaluation frameworks to measure reasoning depth, interaction quality, reliability, and Construct benchmarks reflecting real‑world complexity; establish standards for new architectures Explore adversarial robustness testing, longitudinal tracking, and human‑in‑the‑loop assessment. 🎯 Requirements Experience designing and running evaluations; built or maintained benchmarks or experimental Statistical and analytical rigor; able to extract signal from noisy results. Experience building with models (not just training); proficiency with compound AI systems and Proven track record of research results (publications, notable work). Proactive use of AI tools (e.g., ChatGPT, Cursor, Perplexity) to accelerate workflows. Strong programming and data‑analysis skills; able to prototype ideas and demonstrate effectiveness 🎁 Benefits Base salary range of $150K – $250K, plus meaningful equity and comprehensive benefits. 100% coverage of medical, dental, and vision for employees and dependents. Flexible time off; retirement and financial planning benefits (HSA, FSA, 401(k), etc.). Wellness benefits, including fertility and family-building support. In-office meals and snacks; access to state‑of‑the‑art AI models and tools. Ownership of high‑impact projects across top enterprises; mission‑driven culture.