Job Description
📋 Description Profile ML workloads to identify bottlenecks and optimize training workloads Improve MFU and throughput via parallelism, model compilation, and mixed precision Build observability tools to track MFU, throughput, and latency Develop benchmarking tools to log efficiency gains or regressions Collaborate with Research teams to scale training efficiency 🎯 Requirements 10+ years in performance engineering for ML systems or related fields Experience optimizing large-scale GPU compute jobs Experience with platform teams and research collaborations Proven ability to report and track performance benchmarks Strong Python programming skills BS or MS in ML, CS, Eng, or related field 🎁 Benefits Hybrid work in Sunnyvale, CA Competitive equity package Salary range: $336,400 - $359,000 (USD)