Job Description
📋 Description High-performance inference platform for Grok serving millions of users daily. Design and optimize large-scale model serving end-to-end from distributed infra to low-level GPU Impact-driven role affecting how fast and reliably users interact with Grok at scale. 🎯 Requirements Deep low-level systems programming (C/C++ or Rust) Experience with large-scale, high-concurrent production serving Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.) Strong background in system optimizations: batching, caching, load balancing, parallelism Low-level inference optimizations: GPU kernels, code generation Algorithmic inference optimizations: quantization, speculative decoding, distillation 🎁 Benefits $180,000 - $440,000 USD Equity, comprehensive medical, vision, dental coverage 401(k) retirement plan, disability and life insurance Various discounts and perks