Palo AltoFull TimeEngineering
Remotely
gopythonrustdockerkubernetesgrpcgpu
Job Description
📋 Description
- Develop high-throughput inference systems for internal AI models across SpaceX
- Architect scalable distributed model serving infrastructure with load balancing, auto-scaling
- Optimize model latency/throughput with GPU kernels, quantization, and decoding techniques
- Build reliable, high-concurrency serving systems with low latency and strong observability
- Own end-to-end components: routing, SDKs, rate limiting, and scaling for internal AI inference
- Benchmark, fine-tune, and accelerate inference engines (e.g., SGLang, vLLM, TensorRT-LLM)
🎯 Requirements
- Bachelor's in CS/engineering/math or 2+ years of software experience
- Experience designing and maintaining scalable distributed systems
- 1+ years of full-stack/backend development with production systems
- 1+ years coding in Rust or C++
🎁 Benefits
- Competitive compensation with Level 1-2 ranges
- Comprehensive medical, vision, dental coverage
- 401(k) plan and disability insurance
- Paid parental leave and vacation/time-off
- Sales of SpaceX equity options and employee stock purchase plan
- 10+ paid holidays and paid sick leave