Senior Data Scientist, AI Scoring & Evaluation
WorkeraRemotely
pythonmachine learningsqlllmprompt engineeringstatisticsaidatagpt
Job Description
📋 Description
- Own scoring quality end to end: evaluator design, rubric anchoring, calibration, and accountability
- Build and run the continuous evaluation harness: gold sets, bias diagnostics, drift detection, and
- Define and publish assessment quality KPIs (human to AI agreement, reliability, classification
- Lead the analytics investigations behind assessment decisions: performance studies, impact
- Ship measurement improvements end to end: implement and coordinate with tech, own the spec, the
- Be Workera's liaison between AI governance and InfoSec: stay current on frameworks like GDPR and
🎯 Requirements
- You've driven production data science work independently.
- 4+ years (or 3+ with demonstrated end-to-end ownership) of a production data project ideally in an
- Python and SQL are assumed; you shouldn't need an engineer to run an analysis.
- Real statistical and analytical depth. You can turn, 'is our scoring reliable?' into a defined
- You can read this kind of data. Quantitative work with educational, learning, or assessment data or
- You've worked with LLM-based systems in production. Prompt design, evaluating model output against
Back to all jobs