Senior Principal AI Engineer
JobgetherJob Description
📋 Description Lead design and optimization of large-scale distributed AI training systems across GPU Design, build, and operate distributed training for large neural networks (autoregressive Optimize multi-node, multi-GPU execution for throughput and training efficiency. Diagnose performance bottlenecks across compute, memory, storage, and networking. Improve reliability with fault-tolerant architectures and recovery strategies. Collaborate with research and ML teams to productionize training pipelines. 🎯 Requirements Extensive experience building/operating distributed systems or ML infra. Experience running large-scale GPU workloads in production. Strong PyTorch Distributed experience. Deep understanding of data/tensor/pipeline parallelism. Knowledge of GPU networking (NCCL, RDMA, InfiniBand, NVLink). Experience with Slurm, Kubernetes, Ray, RunAI. 🎁 Benefits Competitive compensation based on experience. Fully remote work opportunity within the United States. Opportunity to work on cutting-edge AI infra and large-scale systems. High-impact role shaping AI development. Collaborative environment with engineers and leaders. Chance to solve challenging distributed computing problems.