San FranciscoFull TimeEngineering
Remotely
pythonrustdistributed systemsllm inference
Job Description
📋 Description
- build scheduling, continuous batching, memory management, KV-cache management, and execution
- develop distributed execution strategies across chips, hosts, and racks, including model
- optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse
- partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove
- enable new model features, execution patterns, numerical formats, and hardware capabilities in a
🎯 Requirements
- strong systems programming experience in C++, Rust, Python, or comparable performance-oriented
- built or optimized runtimes, distributed systems, compilers, kernels, model-serving infrastructure
- understand modern LLM inference, including prefill and decode behavior, batching, KV-cache
- ability to reason quantitatively about latency, throughput, compute intensity, memory bandwidth
- experience profiling and debugging performance across multiple layers of a hardware-software stack
- ability to design clean abstractions while retaining low-level control to extract performance from
🎁 Benefits
- OpenAI standard benefits
- Opportunity to impact frontier AI hardware and software stacks
Back to all jobs