Inference Performance Engineer
MaterialJob Description
📋 Description Build and improve the inference runtime Design scheduling, batching, KV cache, and prefill Implement low-precision kernels and speculative decoding Drive throughput, latency, and cost per token Collaborate with hardware teams on kernels and graph optimizations Own the OpenAI-compatible API surface and serving protocol Build benchmarking and profiling infrastructure 🎯 Requirements BS in CS, EE, or related field or equivalent experience Software: Rust, Go, Python, or C++ Understanding of concurrency, memory, and tail latency Inference tech: transformers, attention, KV cache, batching, quantization Experience with model serving frameworks: vLLM, TGI, SGLang, TensorRT-LLM, llama.cpp GPU/ASIC programming: CUDA, ROCm, Triton Experience with low-precision inference: FP8, FP4, INT4 Profiling and benchmarking with Nsight or perf 🎁 Benefits Top-tier compensation structured to recognize and retain the best talent Meaningful equity Comprehensive medical, dental, vision, life, and disability insurance Parental leave for all new parents, including adoptive and surrogate journeys Flexible PTO Paid Holidays Relocation support