Senior Inference Engineer
JobgetherRemotely
pythondockerkubernetesgolangpytorchvllmtensorrt llmsglang
Job Description
📋 Description
- Build and deploy production-grade LLM inference systems across GPUs
- Design and operate model-serving infra with vLLM, SGLang, TensorRT-LLM
- Optimize workloads for latency, throughput, and costs
- Apply quantization, batching, caching, and routing techniques
- Develop robust infra using Python or Golang with scalable code
- Own initial inference platform with CTO collaboration
🎯 Requirements
- Significant experience building/operating production infra
- Production deployment of LLMs (vLLM, SGLang, TensorRT-LLM)
- Expertise in quantization, batching, caching, and routing
- Strong Python or Golang coding skills
- Understanding of production inference architectures
- Strong problem solving and independent investigation skills
🎁 Benefits
- Competitive compensation including equity
- Health benefits and dependents coverage where available
- Country-specific benefits; remote-friendly schedule
- Remote-first, globally distributed team
- Ownership over architecture and roadmap
- Direct collaboration with senior leadership and Product teams
Back to all jobs