Senior Inference Engineer
JobgetherRemotely
pythondockerkubernetesgolangpytorchvllmtensorrt llmsglang
Job Description
📋 Description
- Build and deploy production-grade LLM inference systems across GPU machines
- Design and operate model-serving infrastructure with vLLM, SGLang, TensorRT-LLM
- Optimize workloads for latency, throughput, reliability, and cost
- Apply quantization, batching, caching, and routing to boost inference performance
- Develop robust production infra using Python or Golang
- Collaborate with CTO and Product to evolve the inference platform and roadmap
🎯 Requirements
- Significant experience building/operating production software or infra systems
- Production deployment of large language models with vLLM, SGLang, TensorRT-LLM or similar
- Expertise in quantization, batching, caching, routing
- Strong Python or Golang production-ready programming skills
- Understanding of production inference architectures and end-to-end user request flow
- Excellent problem-solving and communication skills
🎁 Benefits
- Equity and competitive compensation
- Health/dental/vision/life insurance with dependents coverage where available
- Flexible, outcome-oriented schedule
- Remote-first with globally distributed team
- Ownership over architecture and long-term roadmap
- Collaboration with senior leadership and Product teams
Back to all jobs