North AmericaFull TimeEngineering
Remotely
pythonllmmodel compressionprofilinggpudeepspeeddistributed training
Job Description
📋 Description
- Optimize training and inference workloads to maximize throughput, minimize latency, and improve
- Analyze and improve performance across the full technology stack, including GPU kernels, memory
- Profile CPU, GPU, and distributed workloads to identify bottlenecks and use quantitative analysis
- Design and implement performance improvements using Python, C++, and relevant AI systems
- Optimize distributed training and inference architectures, including model parallelism
- Evaluate and implement model compression techniques while carefully considering their impact on
🎯 Requirements
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related technical
- 6+ years of professional experience in performance engineering, machine learning systems
- Strong programming proficiency in Python and C++, with the ability to develop production-quality
- Hands-on experience optimizing deep learning workloads on modern GPU architectures.
- Deep understanding of distributed training and inference techniques, including parallelism
- Experience using profiling and instrumentation tools across CPU, GPU, and distributed environments.
🎁 Benefits
- $100,000 annual salary for this full-time direct W2 position.
- 100% remote work within the United States.
- Opportunity to work on challenging AI optimization and high-performance computing problems.
- Exposure to large-scale neural networks, modern GPU architectures, distributed systems, and
Back to all jobs