San FranciscoFull TimeEngineering
Remotely
awsllvmgputrainiumcompilersmlirxlakernels
Job Description
📋 Description
- Build and optimize OpenAI's inference stack for AWS Trainium
- Develop high-performance kernels for model ops
- Extend and improve compiler support for Trainium
- Execute and optimize the model forward pass on Trainium
- Profile workloads and identify bottlenecks across stack
- Collaborate with inference/ML systems teams on new models
🎯 Requirements
- 3+ years in ML systems, compilers, kernels, runtimes, or perf engineering
- Strong systems programming fundamentals
- Experience with GPU/TPU/Trainium or similar accelerators
- Ability to reason about perf across hardware to ML frameworks
- End-to-end ownership of complex performance problems
- Bonus: AWS Trainium or Neuron SDK
🎁 Benefits
- Equity
- Competitive salary
- OpenAI benefits package
Back to all jobs