Member of Technical Staff, Performance & Capacity
PSIBostonFull TimeEngineering
Remotely
traininginfinibandgpurdmah100multi node
Job Description
📋 Description
- Measure the fleet with real workloads to guide capacity decisions
- Own and defend the capacity model with clear assumptions
- Manage workload mix and preemptible fractions over time
- Ensure data path speed with parallel file systems and data locality for GPUs
🎯 Requirements
- 5+ years with GPU and large-scale compute workloads
- System-level GPU performance knowledge for heterogeneous GPUs
- Deep understanding of AI training and inference performance
- Experience owning capacity decisions tied to money
- Familiarity with parallel file systems and high-speed networking (InfiniBand, RDMA)
🎁 Benefits
- Competitive compensation and meaningful equity
- Opportunity to work at AI-native infrastructure scale
- Equal opportunity employer with diverse perspectives