Senior Inference Engineer (Remote from Poland)
JobgetherRemotely
pythondockerkubernetesgolangpytorchvllmtensorrt llmsglang
Job Description
📋 Description
- Build and deploy production-grade LLM inference systems across GPU machines.
- Design and operate model-serving infrastructure (vLLM, SGLang, TensorRT-LLM).
- Optimize inference workloads for latency, throughput, reliability, and cost.
- Apply quantization, batching, caching, and routing techniques to improve performance.
- Develop production infrastructure using Python or Golang with maintainable code.
- Own the inference platform evolution with CTO collaboration.
🎯 Requirements
- Significant experience building/operating production software or infra systems.
- Experience deploying/serving LLMs in production (vLLM, SGLang, TensorRT-LLM or similar).
- Experience optimizing inference workloads via quantization, batching, caching, routing.
- Strong Python or Golang production coding experience.
- Understanding of production inference architectures from user request to served response.
- Excellent problem-solving and independent debugging skills.
🎁 Benefits
- Competitive compensation including equity.
- Health, dental, vision, life insurance, with dependents where available.
- Country-specific benefits where applicable.
- Flexible, outcomes-focused schedule.
- Remote-first environment with distributed team.
- Ownership of architecture and long-term roadmap of the inference platform.