Senior Research Data Engineer
CanvaRemotely
pythonmachine learningawsdata pipelinesrlhfmultimodalrayvlm
Job Description
📋 Description
- Design and build data pipelines for agent training across text, image, and multimodal sources
- Build infra for data loading, storage, and retrieval at scale (S3, distributed systems)
- Collaborate with researchers to translate needs into data specs
- Create evaluation datasets and benchmarks with researchers
- Develop tooling for dataset construction, including annotation workflows
- Own data quality, validation, drift monitoring, and reproducibility
🎯 Requirements
- Strong Python software engineering skills
- Production-grade data pipelines and ML DevOps experience
- Experience with ML data workflows, large-scale processing, data versioning
- Hands-on data pipelines for distributed ML training
- Familiar with annotation tooling and human-in-the-loop data collection
- Understanding of ML training needs for good data and downstream impacts
Back to all jobs