Job Description
📋 Description Join the multimodal team to push toward superhuman multimodal intelligence. Work across modalities (image, video, audio, text) from data to end-product experiences. Collaborate with cross-functional teams on frontier capabilities in multimodal reasoning and tool Build models that can see, hear, reason about, and interact with the world in real time. 🎯 Requirements Hands-on experience with multimodal pre-training, post-training, or fine-tuning (vision, audio Expert-level Python, with proficiency in at least one: JAX, PyTorch, or XLA. Experience building or optimizing large-scale distributed ML systems (training/inference Strong data pipelines experience for curation, filtering, generation, and scaling for Strong fundamentals in evaluation design, benchmarks, reward modeling, or RL techniques for Proactive self-starter who thrives in high-intensity environments and owns end-to-end initiatives. 🎁 Benefits Compensation: $180,000 - $440,000 USD. Equity, medical/dental/vision coverage, 401(k), disability and life insurance. Various discounts and perks as part of the total rewards package.