Senior Forward Deployed Engineer I (AI Inference)
DigitalOceanJob Description
📋 Description Embed with AI startups to deploy high-throughput LLM serving Own end-to-end design of multi-tenant AI inference systems Profile bottlenecks and optimize TTFT/TPOT in GPU clusters Collaborate with customers and internal AI infrastructure teams Architect distributed inference with Kubernetes-native tools Deliver production-grade, low-latency serving solutions 🎯 Requirements 6+ years in AI/ML systems or forward deployed engineering Hands-on with vLLM, llm-d, SGLang, TensorRT-LLM Proficient in Python or GoLang; know gRPC; Kubernetes Experience with distributed inference, caching, and batching Strong customer-facing communication and ownership Willing to travel up to 30% 🎁 Benefits Competitive compensation and equity Career development resources and external training Global benefits and flexible time off Equal opportunity employer commitment