Job Description
📋 Description Build state-of-the-art speech systems end-to-end—from data specs through production inference. Drive the model ↔ data ↔ eval flywheel for VC and related tasks (controllable TTS, voice design). Collaborate with research, data, and infra to ship fast, reliable, cost-aware models. Own work end-to-end, bridging research and engineering. Design experiments, listen tests, and metrics aligned with user-perceived quality. Develop robust tooling and contribute to the full stack from low-level optimizations to model 🎯 Requirements Exceptional research/development experience with large-scale audio models (>8B parameters Hands-on with diffusion/flow-matching transformers; practical knowledge of samplers, schedules Experience training audio VAEs, neural audio codecs, vocoders; perceptual objectives, adversarial Multi-node, multi-GPU distributed training (FSDP/DeepSpeed or equivalent). Strong software engineering skills; production-quality code. Proficiency with PyTorch; performance work (profiling, CUDA/Triton/C++). 🎁 Benefits Competitive salary with equity Medical, dental, vision insurance Generous PTO and holidays Parental leave & fertility support 401(k) retirement savings plan Lifestyle spending account