Job Description
📋 Description Lead architecture, performance, and reliability for Kubernetes-native inference at massive scale. Define cross-cutting design initiatives: routing, adaptive scheduling, GPU resource mgmt Implement advanced inference optimizations: speculative decoding, KV-cache reuse, benchmarking Aim for strict P99 SLAs with metrics-driven engineering and observability. Collaborate across infrastructure boundaries and mentor senior/mid engineers. 🎯 Requirements 8–12+ years in large-scale distributed systems or cloud platforms. Proven track record leading cross-team technical initiatives at scale. Strong coding in Go, Python, or C++. Deep Kubernetes production-scale experience (orchestration, scheduling, service design). Strong knowledge of networked systems, performance optimization, distributed design. Hands-on inference systems experience: batching, caching, memory opt, mixed precision (BF16/FP8) 🎁 Benefits Competitive salary with discretionary bonus and equity. Comprehensive benefits package including health, dental, vision. 401(k) with employer match and flexible PTO. Casual work environment and growth opportunities at a fast-growing AI/cloud company.