Job Description
📋 Description Design scalable LLM-as-a-judge pipelines to score actions and tool usage. Develop rubrics and golden datasets to baseline agent performance. Optimize prompts for judges to ensure high-quality scoring. Full-stack contribution to core AI features, UI components, and architecture. Analyze and report evaluation metrics to identify failures and regressions. Onsite collaboration Tue/Wed; Thurs encouraged. 🎯 Requirements Hands-on LLM evaluation experience designing and testing automated tests. Strong Python background; experience with TypeScript/Node.js. Prompt engineering expertise for reliable model prompting. Familiarity with LLM evaluation tooling (e.g., LangSmith, Braintrust, promptfoo). Agentic architectures understanding (ReAct, tool-use APIs). Analytical mindset to translate subjective signals into metrics. 🎁 Benefits Competitive compensation with ownership program. Flexible work culture across remote, hybrid, and in-office spaces. Generous time off including holidays and the Dim the Lights period. Wellness programs and mental health support. Learning and development resources and tuition reimbursement. The technology and tools you need to do your best work.