Senior HPC Cluster Engineer - AI, ML
JobgetherRemotely
kubernetesslurmrhellsfcentospbsbcmrtda
Job Description
📋 Description
- Provide technical leadership for systems administration and service delivery across large-scale
- Own the day-to-day operation of production AI/HPC clusters, monitoring system health, user
- Build, maintain, and scale heterogeneous AI/ML clusters across on-premises and cloud environments
- Develop scalable automation and tooling to improve the deployment, configuration, management, and
- Collaborate with global engineering teams and internal users to understand evolving research and
- Support researchers in running complex AI/HPC workloads, including performance analysis
🎯 Requirements
- Bachelor's degree in Computer Science, Electrical Engineering, or a related discipline, or
- 5+ years of experience designing, deploying, and operating large-scale compute or infrastructure
- Strong experience with AI/HPC job schedulers such as Slurm, Kubernetes, PBS, RTDA, BCM, or LSF.
- Proficiency administering Linux distributions such as CentOS/RHEL and/or Ubuntu.
- Hands-on experience with cluster configuration and infrastructure management tools including BCM
- Strong knowledge of container technologies such as Docker, Singularity, Podman, Shifter, or
🎁 Benefits
- Opportunity to work on large-scale AI and high-performance computing infrastructure supporting
- Exposure to GPU-accelerated computing, distributed systems, cloud infrastructure, HPC, and emerging
- Collaboration with highly skilled engineering and research teams across global locations.
- Opportunities to influence infrastructure architecture, automation, reliability, and operational
- Continuous learning and exposure to rapidly evolving technologies in AI, ML, and high-performance
- Flexible work arrangements may be available, with opportunities based in Bengaluru, Pune, or remote
Back to all jobs