Senior AI Infrastructure & Platform Operations Engineer (remote in the EU)
MirantisAlso available on
Job Description
📋 Description Lead technical operations for large-scale AI infra with NVIDIA GPUs, Kubernetes, and Serve as escalation point during critical incidents and drive reliability initiatives. Shape operational standards, automation, and platform evolution with k0rdent AI. Collaborate with data centers, vendors, and engineering teams to resolve complex challenges. Participate in incident management and service restoration activities. 🎯 Requirements 7+ years in infrastructure/platform operations, SRE, or related roles. Expert Linux administration and troubleshooting. Strong networking and production Kubernetes experience. Experience supporting large-scale production infra and distributed systems. Root cause analysis and long-term operational improvement experience. Observability and monitoring expertise; strong communication skills. 🎁 Benefits Opportunities to work with advanced AI infra and NVIDIA GPUs. Collaborate with top engineers on scale challenges.