Senior AI Infrastructure Platform Operations Engineer (remote in the US)
MirantisJob Description
📋 Description Lead investigations and resolution of complex infra, networking, and platform incidents. Senior escalation point for operational events during critical service impacts. Support large-scale NVIDIA GPU infrastructure and high-performance networking. Troubleshoot Linux, Kubernetes, storage, and hardware issues. Analyze platform performance, capacity, stability, and reliability trends. Lead root cause analysis and drive long-term corrective actions. 🎯 Requirements 7+ years of experience in infrastructure operations, platform operations, SRE, or related roles. Expert-level Linux administration and troubleshooting. Strong networking expertise for diagnosing performance and reliability issues. Production experience operating Kubernetes in production environments. Experience leading technical investigations and incident management. Strong observability, monitoring, and reliability practices. 🎁 Benefits Operate advanced AI infrastructure environments in production today. Work with NVIDIA GPUs, Kubernetes, and high-performance networking tech. Help define operational standards and reliability practices for AI infra. Influence the adoption of AI-powered operational capabilities via k0rdent AI. Collaborate with highly skilled engineers solving complex infra challenges at scale. Competitive compensation with strong benefits package.