SunnyvaleFull TimeEngineering
Remotely
pythontensorrttritonncclmegatrondeepspeedonnx runtime
Job Description
📋 Description
- Profile and optimize distributed training end to end
- Optimize offline and batch inference over petabyte logs
- Establish roofline and performance models
- Improve multi-node scaling efficiency
- Drive cluster goodput and reduce GPU idle time
- Build benchmarking, observability, and regression tooling
🎯 Requirements
- Hands-on ML performance engineering experience
- Experience with distributed multi-node training at scale
- Deep familiarity with GPU performance concepts
- Experience with high-throughput or batch inference systems
- Fluency in Python and proficiency in C++
- Strong debugging, analytical, and problem-solving skills
🎁 Benefits
- Collaborative culture and ownership mindset
- Equal opportunity employer commitments