Senior Inference Engineer
JobgetherRemotely
pythondockerkubernetesgolangpytorchvllmtensorrt llmsglang
Job Description
📋 Description
- Build and deploy production-grade LLM inference systems across one or multiple GPU machines, owning
- Design, implement, and operate model-serving infrastructure using technologies such as vLLM
- Optimize inference workloads for scale, balancing latency, throughput, reliability, and
- Apply techniques such as quantization, batching, caching, and intelligent request routing to
- Develop robust production infrastructure using Python or Golang, with an emphasis on maintainable
- Establish the initial inference platform in close collaboration with the CTO and take ownership of
🎯 Requirements
- Significant professional experience building and operating production software or infrastructure
- Demonstrated experience deploying and serving large language models in production, ideally using
- Practical expertise optimizing inference workloads through quantization, batching, caching
- Strong programming skills in Python or Golang, with a track record of writing and maintaining
- Strong understanding of production inference architectures, including the journey from user request
- Excellent problem-solving skills and the ability to independently investigate complex performance
🎁 Benefits
- Competitive compensation package including equity.
- Health, dental, vision, and life insurance, with coverage for eligible dependents where available.
- Benefits adapted to the country of employment.
- Flexible working schedule focused on outcomes rather than fixed working hours.
- High degree of workplace flexibility, supporting remote work and changing personal circumstances.
- Remote-first environment with a globally distributed team.