Research Engineer - Inference
ElevenLabsRemotely
vllmcudatensorrttritongpuinferencesglangserving
Job Description
📋 Description
- Deploy state-of-the-art models to production
- Own path from research checkpoint to serving infra
- Optimize inference latency, throughput, and cost
- Use quantization, distillation, KV-cache, batching, custom kernels
- Build high-performance serving systems for real-time workloads
- Create tooling to ship models quickly and safely
🎯 Requirements
- Experience deploying and serving ML models in production
- GPU programming and inference optimization: CUDA, Triton, TensorRT
- Experience with vLLM, SGLang or similar serving frameworks
- Ability to profile, diagnose, and eliminate bottlenecks
🎁 Benefits
- Innovative culture with impact-focused work
- Growth paths and opportunities to drive impact
- Learning & development stipend
- Social travel stipend for annual meetups
- Annual company offsite in new locations
- Monthly co-working stipend
Back to all jobs