AI Infrastructure Engineer
MirantisJob Description
📋 Description Manage and operate production AI infrastructure environments. Lead incident response and restore service during outages. Troubleshoot infrastructure and networking across bare-metal and cloud. Conduct root cause analysis and drive improvements. Improve operational documentation and knowledge base. Collaborate with global team across time zones for continuous coverage. 🎯 Requirements Proven experience managing and operating large-scale production systems (bare-metal and cloud). Kubernetes knowledge with strong troubleshooting skills. Experience configuring, customizing, and extending logging and monitoring tools (Prometheus, Grafana, ELK). Experience with infrastructure automation technologies and IaC practices (Ansible, Terraform). Effective verbal and written communication in English. Strong analytical and problem-solving skills; willing to work weekends/holidays. 🎁 Benefits Work with an established Silicon Valley leader in cloud infrastructure. Collaborate with passionate colleagues helping Fortune 500 and Global 2000 customers. Be part of cutting-edge, open-source innovation. Thrive in a high-energy, open and collaborative culture. Professional development and training; attend conferences and working groups. Competitive compensation and strong benefits package.