Also available on
Job Description
📋 Description Lead the AI/ML Platform Inference team to productize and operate the offering. Guide engineers on reliability, observability, packaging, tooling, and dev experience. Partner with Product, CoreWeave engineering, Design, Support, and GTM to deliver platform capabilities with impact. Own the inference productization layer: incident response, releases, observability, and quality. Collaborate with infrastructure teams to deliver a cohesive, user-facing product. Drive tracing and tooling enhancements for application-layer workloads. 🎯 Requirements 7+ years in software engineering; 3+ years leading distributed systems teams. Fluent in observability, autoscaling, reliability, API design, and service operations. Balance latency, reliability, cost, and development velocity amid competing priorities. Strong leadership and communication across multiple stakeholders. Empathy for ML practitioners and platform developers; improve reliability and UX. Background in high-scale systems or cloud infrastructure; model-serving/inference a plus. 🎁 Benefits Medical, dental, and vision insurance - 100% paid by CoreWeave 401(k) with generous employer match Tuition Reimbursement Health Savings Account Paid Parental Leave Flexible PTO