Job Description
📋 Description Build scalable, observable APIs for model inference and fine-tuning of LLMs. Collaborate with ML researchers and stay current on industry trends. Participate in on-call to ensure service availability. Own projects end-to-end: requirements, design, to code. Make cost-efficient build vs buy tradeoffs with good automation. 🎯 Requirements 2+ years building ML training pipelines or inference services in production. Experience deploying, fine-tuning, training, and prompting LLMs. Experience with LLM inference latency optimization techniques (quantization, DeepSpeed). Strong ML fundamentals and production-grade infra experience. Proficiency in Python, Docker, Kubernetes, and Terraform. SF or NYC based and able to be in person. 🎁 Benefits Health, dental, and vision coverage; retirement benefits. Learning and development stipend; generous PTO. Commuter stipend and other perks.