Job Description
📋 Description Accelerate Inference: Lead and implement advanced inference acceleration techniques, including Maximize GPU Parallelism: Engineer and optimize GPU strategies across tensor, sequence, and Programming for Performance: Develop and optimize high-performance computing kernels and Advance AI Deployment: Bring state-of-the-art videogen and large language models into production in Improve Training Efficiency (Bonus): Contribute to improvements in model training speed, stability 🎯 Requirements Experience: 5+ years engineering experience, with a strong track record in inference acceleration Inference Mastery: Expertise in inference optimization, including quantization, attention GPU & Parallelism: Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP AI Domain Knowledge: Familiarity with video generation (videogen) models and large language models Collaboration & Ownership: Strong cross-discipline communication, self-driven 🎁 Benefits Competitive salary in the AI industry Equity in a fast-growing startup shaping the future of AI Comprehensive health benefits, monthly stipends, company retreats A supportive and collaborative office culture—we’re all building and launching together