Senior Research Engineer - Voice
SynthesiaAlso available on
Job Description
📋 Description Develop streaming and speech-to-speech systems for low-latency voice synthesis. Adapt models for conditioning inputs (emotion, speed, prosody, speaker control). Implement post-training optimizations (quantization, pruning, distillation) for real-time speech. Test novel architectures (neural codecs, diffusion, flow-matching) for realism. Define metrics for conversational speech, incl. latency-aware MOS. Apply DPO and distillation to fine-tune large-scale speech models. 🎯 Requirements Strong understanding of generative modeling for sequential or multimodal data. Experience with LLMs or transformer architectures. Proficiency in PyTorch, with distributed training and optimization. Time-series modeling and tokenization for audio or speech. Prototype quickly, test hypotheses, iterate efficiently. End-to-end DL model training from data prep to evaluation. Software engineering skills enabling contributions to large research infra.