Senior Inference Engineer
JobgetherRemotely
pythonkubernetesgolangpytorchvllmcudatensorrt llmsglang
Job Description
📋 Description
- Build and deploy production-grade LLM inference systems across GPU machines.
- Design and operate model-serving infra using vLLM, SGLang, TensorRT-LLM.
- Optimize workloads for latency, throughput, and cost.
- Develop robust infra with Python or Golang for production use.
🎯 Requirements
- Extensive experience building/operating production infra.
- Production deployment of LLMs (vLLM, SGLang, TensorRT-LLM, etc.).
- Expertise in quantization, batching, caching, routing techniques.
- Strong Python or Golang production code skills.
- Understanding of production inference architectures.
- Problem-solving and independent investigation skills.
🎁 Benefits
- Competitive compensation including equity.
- Health, dental, vision, life insurance; dependents coverage where available.
- Benefits adapted to country of employment.
- Remote-friendly, flexible schedule and global team.
- Ownership of architecture and roadmap for inference platform.
- Collaboration with senior leadership and product teams.
Back to all jobs