Forward Deployed Engineer (Inference & Post-Training) - Mandarin Speaking
Together AIRemotely
pythonvllmsfttensorrt llmloradposglanginference engine
Job Description
📋 Description Hands-on technical partner to strategic customers Focus on inference, post-training, and optimization Collaborate with SAs on POCs and deployments Drive time-to-value with onboarding and configurations Influence product roadmap from field insights 🎯 Requirements 5+ years in a technical role Expert with inference engines (e.g., vLLM, TensorRT-LLM) Deep KV cache, decoding, parallelism, quantization expertise Post-training experience: LoRA, SFT, DPO, RLHF, GRPO Strong Python in production settings Open-source model awareness for customer use cases 🎁 Benefits Competitive compensation and startup equity Health insurance and other benefits Flexibility in remote work
Back to all jobs