Job Description
📋 Description Maintain and scale research infrastructure to maximize system performance. Collaborate with research teams to design cost-efficient infrastructure solutions. Identify bottlenecks and optimize performance across distributed systems. Build telemetry and monitoring for cloud and datacenter fleets. Participate in on-call rotations to ensure reliability. 🎯 Requirements Experience building or operating large-scale training platforms Experience with large-scale compute clusters (GPUs) Debug performance and reliability across distributed fleets Strong problem-solving and independent working ability Strong communication with internal and external partners Deep knowledge of cloud infra: Kubernetes, IaC, AWS, GCP 🎁 Benefits Hybrid work model with offices in Freiburg/SF and periodic travel Competitive salary with equity Relocation encouraged but not required Travel costs covered for in-person collaboration