North AmericaFull TimeEngineering
Remotely
gopythonansibleterraformprometheusclickhousegrafananccl
Job Description
📋 Description
- Build and operate monitoring and alerting for cluster health — fabric, GPU, power/thermal, and
- Remotely deploy and configure large-scale HPC clusters for AI workloads using automation
- Automate cluster lifecycle: OS, firmware, drivers, networking as code (Ansible, Terraform)
- Create runbooks and automated remediations for common cluster failure modes
- Troubleshoot and resolve cluster issues across InfiniBand/RoCE, NCCL, GPU-direct, fabric
- Participate in on-call rotations and lead incident response for cluster-level problems
🎯 Requirements
- 7+ years of experience in Site Reliability Engineering, HPC Engineering, DevOps, or similar
- Strong understanding of modern AI infrastructure including GPU architectures
- Solid Linux-based systems knowledge in distributed environments
- Experience with InfiniBand (IB), RoCE, CLOS fabrics, 100GbE, NCCL
- Proficiency in Python and Go; experience improving internal tooling
- Experience with monitoring and alerting tools (Prometheus, Grafana, Clickhouse); automation with
🎁 Benefits
- Generous cash and equity compensation
- Health, dental, and vision coverage for you and dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off
- Equal Opportunity Employer