Senior Inference Engineer
JobgetherRemotely
pythondockerkubernetesgolangpytorchvllmtransformerstensorrt llm
Job Description
📋 Description
- Build and deploy production-grade LLM inference systems across GPU machines
- Design and operate model-serving infra using vLLM, SGLang, TensorRT-LLM
- Optimize workloads for latency, throughput, reliability, and cost
- Apply quantization, batching, caching, and routing techniques
- Develop production code in Python or Golang with emphasis on scalability
- Own initial inference platform and evolve it as the org scales
🎯 Requirements
- Significant experience building/operating production software or infra systems
- Production deployment of large language models, ideally with vLLM, SGLang, TensorRT-LLM
- Experts in quantization, batching, caching, routing for inference
- Strong Python or Golang coding skills for production-quality code
- Understanding of production inference architectures from user request to served response
- Excellent problem-solving and independent investigation skills
🎁 Benefits
- Competitive compensation with equity
- Health, dental, vision, life insurance, with dependents coverage where available
- Benefits adapted to country of employment
- Flexible, outcome-focused schedule; remote-friendly
- Remote-first environment with distributed team
- Ownership over architecture, implementation, and long-term roadmap
Back to all jobs