Job Description
📋 Description Turn research checkpoints into production-ready inference services Design and maintain high-performance APIs serving millions of requests Optimize inference latency and throughput across GPU infrastructure Build scalable serving architectures that handle unpredictable traffic Improve reliability, monitoring, and observability across model-serving systems Prototype and ship demos that showcase new capabilities in days, not weeks 🎯 Requirements Building and operating ML inference services in production Designing scalable API architectures with async processing Optimizing GPU workloads (batching, quantization, CUDA) Managing distributed systems and task queues under variable load Implementing monitoring and observability for production ML systems Debugging performance bottlenecks across model, infrastructure, and network layers 🎁 Benefits Distributed team with offices in Freiburg and SF, with options for remote work and periodic Travel costs covered to facilitate in-person collaboration Strong engineering culture focused on research excellence and collaboration Opportunity to work at the intersection of backend systems, GPU performance, and production ML