HPC Infrastructure Site Reliability Engineer
RadiantJob Description
📋 Description Operate and improve high-density AI/HPC infrastructure in a 24/7 production environment Participate in a 24x7x365 on-call rotation, incident response Troubleshoot complex issues across compute, networking, storage, and orchestration layers in GPU-accelerated environments Lead performance evaluation and acceptance of new HPC infrastructure before production Drive continuous service improvement through automation, tooling, and process refinement Build and maintain infrastructure automation and tooling (IaC and scripting) 🎯 Requirements 8+ years in SRE/Infra in large-scale 24/7 production 2–3+ years HPC/AI infra with GPU compute at scale Strong Linux (Ubuntu) admin and production troubleshooting Bare-metal infra and out-of-band tooling (IPMI/iLO/iDRAC/Redfish) Solid networking incl. InfiniBand and RoCE Automation/scripting (Bash, Python, Ansible) and IaC 🎁 Benefits Work with cutting-edge GPU/AI infrastructure Global, distributed team with learning culture Flexible, globally connected workplace Opportunity to grow with a scaling business