San FranciscoFull TimeEngineering
Remotely
machine learningvideodeep learningaudiopytorchtensorflowmultimodalvision language
Job Description
📋 Description
- Generative media synthesis for video and audio features (production research).
- Zero-shot voice and roomtone cloning from minutes of audio.
- Multimodal understanding with vision-language systems for editing features.
- Computer-vision-heavy problems like digital human reconstruction and facial modeling.
- Develop new algorithms for media synthesis, speech, anomaly detection, and tagging.
- Identify and pursue next research directions to become features, not just papers.
🎯 Requirements
- Proven ability to design and implement deep learning algorithms with publications, OSS, or shipped
- Strong PyTorch and/or TensorFlow skills.
- Track record of generating new ML ideas and running experiments quickly.
- Strong experimental judgment and honesty about what works.
- PhD or Master's in deep learning, or equivalent experience.
🎁 Benefits
- Base salary range: $261,625–$299,000.
- Equity and benefits; location- and level-dependent offers.
Back to all jobs