Senior Inference Engineer
JobgetherRemotely
pythondockerkubernetesgolangpytorchvllmtensorrt llmsglang
Job Description
📋 Description
- Build and deploy production-grade LLM inference systems across GPU machines.
- Design and operate model-serving infra with vLLM, SGLang, TensorRT-LLM.
- Optimize latency, throughput, reliability, and cost.
- Apply quantization, batching, caching, and routing for performance.
- Develop infra in Python or Golang with scalable, maintainable code.
- Own initial inference platform with CTO and evolve it as the org scales.
🎯 Requirements
- Extensive experience building and operating production software or infra systems.
- Experience deploying LLMs in production (vLLM, SGLang, TensorRT-LLM or similar).
- Expertise optimizing workloads via quantization, batching, routing, etc.
- Strong Python or Golang production-grade coding skills.
- Understanding of production inference architectures from request to response.
- Excellent problem-solving and independent diagnostic ability.
🎁 Benefits
- Competitive package including equity.
- Health, dental, vision, life insurance; dependents coverage where available.
- Country-specific benefits.
- Flexible, outcome-driven schedule and remote-first environment.
- Global distributed team and significant ownership over architecture.
- Direct collaboration with senior leadership and Product teams.
Back to all jobs