San FranciscoFull TimeEngineering
Remotely
kubernetesroceinfinibandrdmadeepspeedmegatron lmpytorch fsdpnvidia a100/h100/h200
Job Description
📋 Description
- Identify new market opportunities for Core Services compute and networking capabilities.
- Drive first wins by translating field insight into compute and networking roadmap inputs.
- Lead front-edge GTM for Compute (bare-metal/virtual GPUs) and Networking (high-speed interconnects).
- Turn customer requirements into product feedback to shape roadmap.
🎯 Requirements
- 10+ years in HPC, data center infra, or GPU cluster engineering with revenue outcomes.
- 5+ years with large-scale GPU clusters, RDMA networking, and distributed training.
- Deep knowledge of high-speed fabrics (InfiniBand, RoCE, NVLink) and topology effects on scale.
- Experience deploying large Kubernetes clusters for GPU workloads (NVIDIA Operator, MIG
- Strong understanding of compute types, memory bandwidth, NVMe storage, power/cooling for workloads.
- Familiarity with PyTorch FSDP, Megatron-LM, DeepSpeed, and multi-region deployments.
🎁 Benefits
- Competitive base salary with discretionary bonus and equity.
- Comprehensive medical/dental/vision and company-paid life insurance.
- Flexible PTO, 401(k) with employer match, and ESPP.
- Health and family support programs, including tuition reimbursement and childcare assistance.
- Collaborative, fast-paced culture with opportunities for growth.