San Francisco, SunnyvaleFull TimeEngineering
Remotely
jiralinuxincident managementhpcpcienvlinkgpunvidia
Job Description
📋 Description
- Lead daily support operations for Bare Metal infrastructure across client environments.
- Triaging incidents, drive escalations, and ensure hardware is monitored and delivered effectively.
- Build and lead Infrastructure Support team; collaborate with product, engineering, and infra teams.
- Manage client communications during escalations to ensure transparency and satisfaction.
- Mentor team members and scale processes as we grow.
🎯 Requirements
- 5+ years leading infrastructure support, data center ops, or physical compute teams.
- Hands-on Linux administration and CLI proficiency.
- Experience with hardware diagnostics, troubleshooting, and replacements (servers, power, cabling).
- Experience with high-performance rack-scale hardware and GPU compute nodes.
- Understanding of GPU infrastructure (NVIDIA A100/H100s, PCIe/NVLink, liquid cooling) or quick
- Proven incident/escalation management ownership of production-impacting issues.
🎁 Benefits
- Competitive benefits package; medical, dental, vision insurance.
- Company-paid life insurance; disability, FSA/HSA, 401(k) with employer match.
- Paid Parental Leave; flexible PTO; tuition reimbursement.
- Employee Stock Purchase Program (ESPP); AI/ML workload-friendly culture.
Back to all jobs