Remotely
trainingpytorchcudatritonlow latencygpudiffusionautoregressive
Job Description
📋 Description
- Train and optimize large-scale video and multimodal models
- Improve efficiency across training and inference (memory, latency, cost)
- Implement techniques such as distillation, quantization, and pruning to aggressively accelerate
- Build and maintain distributed training systems
- Optimize GPU utilization, parallelism, and throughput
- Translate research models into robust, production-ready systems
🎯 Requirements
- BS/MS/PhD in CS, ML, or related field
- 2+ years of professional industry experience
- Strong experience in deep learning systems and infrastructure
- Expertise in PyTorch, CUDA, Triton, and distributed training (FSDP, etc.)
- Experience scaling and optimizing large models under low-latency inference constraints
- Strong debugging and performance profiling skills
🎁 Benefits
- Comprehensive medical, dental, and vision plans
- 401K with employer match
- Commuter Benefits
- Catered lunch multiple days per week
- Dinner stipend every night if you're working late and want a bite!
- Grubhub subscription
Back to all jobs