Job Description
📋 Description Own infrastructure to run training and inference workloads on GPU clusters. Build a self-serve platform for inference engineers and researchers. Operate the GPU fleet across providers with unified provisioning. Develop scheduling, placement, and fault tolerance for GPUs. Manage Kubernetes for GPU orchestration and multi-cluster ops. Ensure reliability and observability of the platform. 🎯 Requirements Deep Kubernetes experience with custom operators and CRDs. Experience managing GPU clusters, NVIDIA hardware, CUDA, networking. Experience across multiple clouds (CoreWeave, AWS, GCP or similar). Strong distributed systems fundamentals: scheduling, resource allocation, fault tolerance. Proficiency in Go, Rust or C++ for infra and systems code.