Job Description
📋 Description Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training Manage and optimize Slurm-based HPC environments for distributed training of large language models Develop robust APIs and orchestration systems for both training pipelines and inference services Implement resource scheduling and job management systems across heterogeneous compute environments Benchmark system performance, diagnose bottlenecks, and implement improvements across both training Build monitoring, alerting, and observability solutions tailored to ML workloads running on 🎯 Requirements Expert-level Kubernetes administration and YAML configuration management Proficiency with Slurm job scheduling, resource management, and cluster configuration Experience deploying and managing distributed training systems at scale (PyTorch) Deep understanding of container orchestration and distributed systems architecture Familiarity with LLM concepts and training processes; GPU resource optimization Experience operating large-scale Kubernetes deployments in production