Senior AI Infrastructure Platform Operations Engineer (remote in the EU)
MirantisJob Description
📋 Description Lead technical operations for large-scale AI infrastructure with NVIDIA GPUs, Kubernetes, and Escalation point for incidents and critical issues in production environments. Shape reliability practices, automation initiatives, and platform evolution for AI-powered services. Oversee platform performance, capacity, and reliability trends; drive long-term improvements. Collaborate with engineering, hardware vendors, and datacenters to resolve challenges. Participate in major incident management and service restoration. 🎯 Requirements 7+ years in infrastructure/platform operations, SRE, or related roles. Expert Linux administration and troubleshooting. Strong networking and production Kubernetes experience. Experience supporting large-scale distributed systems and incident leadership. Proven root cause analysis and operational improvement skills. Experience with observability/monitoring and reliability practices. 🎁 Benefits Operate advanced AI infrastructure with latest NVIDIA GPUs and high-performance networking. Influence operational standards and reliability practices for next-gen AI platforms. Opportunity to contribute to k0rdent AI-powered capabilities. Collaborate with skilled engineers on complex scale challenges.