Mountain ViewFull TimeEngineering
Remotely
gopythonkubernetesdistributed systemsncclgpu
Job Description
📋 Description
- Work on ML infrastructure that trains models at the core of the Nuro Driver™
- Own training pipelines, multi-generation accelerators, and multi-cluster scheduling/orchestration
- Design and operate large-scale data pipelines (batch/streaming, storage layout, high throughput)
- Develop agentic-first workflows (data-to-training-to-evaluation) that are reproducible and
- Ensure reliability for training and release pipelines; implement monitoring, alerts, and incident
🎯 Requirements
- BS, MS, or PhD in CS/EE or related field; 3+ years of relevant work experience
- Willingness to deep-dive into implementation and raise tech/operational standards
- Ownership mindset with focus on monitoring, alerting, runbooks
- Strong proficiency in Python; comfort with C++, Go or similar systems language
- Hands-on experience running production infrastructure on Kubernetes
- Solid distributed-systems fundamentals, performance/reliability reasoning
🎁 Benefits
- Base pay range: $193,930 - $291,150 (USD), eligible for annual bonus, equity, and benefits