Product Manager - AI Inference & Model Serving
MirantisJob Description
📋 Description Own strategy, roadmap, and lifecycle for inference and model serving Define serverless inference, dedicated endpoints, autoscaling, and routing Manage KV cache and observability for production inference workloads Collaborate with NeoClouds and enterprise teams on requirements and architecture Align latency, throughput, utilization, and cost with measurable outcomes 🎯 Requirements 7+ years in product management or senior technical AI/ML roles Strong knowledge of production AI inference: model serving, autoscaling, observability Able to reason about GPU, network, storage, orchestration trade-offs Experience with runtimes: vLLM, SGLang, TensorRT-LLM, Dynamo, Triton Comfortable in architecture reviews and technical-commercial discussions with platform teams 🎁 Benefits Work with an established Silicon Valley leader in cloud infrastructure Collaborate with passionate colleagues across Fortune 500 and Global 2000 customers Be part of cutting-edge, open-source innovation Thrive in a high-energy environment with openness, collaboration, risk-taking, and growth Professional development and training Attend conferences and working groups