North AmericaFull TimeEngineering
Remotely
kubernetessparkairflowflinkraykubeflowargo
Job Description
📋 Description
- Design and build large-scale offline ML experimentation platforms for reproducible research, model
- Develop production-grade training orchestration frameworks for distributed training, hyperparameter
- Build infrastructure for experiment tracking, metadata management, lineage, artifact versioning
- Collaborate with ML engineers and researchers to improve experimentation velocity and operational
- Create automated workflows for model promotion, rollback, compliance validation, and continuous
- Design an agentic AI execution platform supporting autonomous and human-in-the-loop workflows
🎯 Requirements
- 5+ years in infrastructure/platform engineering or large-scale distributed systems.
- 2+ years building/operating production ML infrastructure, SDKs, APIs, or self-service AI tooling.
- Experience with distributed data processing (e.g., Spark, Flink, Ray).
- Experience with orchestration/workflow tech (Kubeflow, Argo, Airflow).
- Experience building offline ML experimentation platforms, registries, tracking systems, or training
- Experience with agentic AI systems, multi-agent orchestration, autonomous workflows, and related
🎁 Benefits
- Comprehensive Healthcare Benefits and Income Replacement Programs
- 401k with Employer Match
- Global Benefit programs for workspace, development, and caregiving
- Family Planning Support
- Gender-Affirming Care
- Mental Health & Coaching Benefits
Back to all jobs