Job Description
📋 Description Serve as a senior escalation point, troubleshooting hardware/driver/kernel issues Quickly distinguish hardware failures, driver issues, kernel problems, and misconfig Proactively identify process/tooling/documentation gaps and fix them Use AI tools to build scripts and small internal tools to close gaps Perform root-cause analysis across distributed systems, clusters, and GPU infrastructure Collaborate with engineering to turn recurring pain points into permanent fixes 🎯 Requirements 3+ years of hands-on HPC experience in admin, support, or engineering Very strong Linux system administration experience Experience with HPC environments; Linux cluster admin; Kubernetes/Slurm Strong coding and CI/CD experience; AI-assisted tooling Proficiency with monitoring/logging tools: Prometheus, Grafana, Datadog CUDA/NCCL/NVLink/GPUDirect RDMA; IB/RoCE networking 🎁 Benefits Generous cash and equity compensation Health, dental, and vision coverage for you and dependents Wellness and commuter stipends for select roles 401k plan with 2% company match (USA employees) Flexible paid time off plan Equal Opportunity Employer