Senior Inference Engineer
JobgetherRemotely
pythondockerkubernetesgolangpytorchvllmtensorrt llmsglang
Job Description
📋 Description
- Build and deploy production-grade LLM inference systems across GPU machines
- Design and operate model-serving infra using vLLM, SGLang, TensorRT-LLM
- Optimize inference workloads for latency, throughput, reliability, cost
- Apply quantization, batching, caching, routing to boost perf
- Develop infra in Python or Golang with scalable engineering focus
- Partner with CTO to shape initial inference platform and roadmap
🎯 Requirements
- Significant experience building/operating production software or infra
- Production deployment of large language models (vLLM, SGLang, TensorRT-LLM, or equivalent)
- Optimization techniques: quantization, batching, caching, routing
- Strong Python or Golang production code skills
- Understand end-to-end inference architecture from request to response
- Excellent problem-solving and independent investigation skills
🎁 Benefits
- Competitive package with equity
- Health/dental/vision/life insurances
- Benefits tailored to country of employment
- Remote-friendly with flexible schedule
- Global distributed team with ownership and impact
- Work on GPU/AI infra at scale and open-source tech
Back to all jobs