Job Description
📋 Description Embed in production training and tackle hard training-system problems Improve performance, reliability, and numerical stability of production runs Profile and optimize full training steps across model, attention, kernels, data loading, and memory Debug distributed training failures and build profiling harnesses Translate architecture changes into efficient training implementations Collaborate with researchers to move research into production-ready code 🎯 Requirements Experience on large-scale training systems, ideally with researchers Strong PyTorch fluency; comfortable with low-level training code Distributed training concepts: FSDP, tensor/model/context/sequence parallelism, NCCL Hands-on experience improving throughput, memory, or stability GPU profiling with Nsight Systems/Compute, torch profiler, trace viewers, or custom telemetry Understanding of low-precision training and quantization tradeoffs (FP8, MXFP8, FP4) 🎁 Benefits Distributed team with offices in Freiburg and SF; in-person weeks and travel support Collaboration across researchers and engineers to advance frontier models