AI Infrastructure Engineer
MirantisRemotely
gopythonkubernetesansibleterraformprometheusgrafanaelk
Job Description
📋 Description
- Manage and operate production AI infrastructure environments.
- Lead incident response and timely service restoration during outages.
- Troubleshoot infrastructure and networking across bare-metal and cloud with multiple vendors.
- Conduct root cause analysis and drive improvements.
- Improve operational docs and knowledge base.
- Collaborate with global team across time zones, including occasional weekends/holidays.
🎯 Requirements
- Experience managing large-scale production systems (bare-metal and/or cloud).
- Strong Kubernetes knowledge with solid troubleshooting.
- Experience configuring/logging/monitoring tools (Prometheus, Grafana, ELK or similar).
- Infrastructure-as-Code: Ansible, Terraform or similar.
- English communication skills, both written and verbal.
- Analytical problem-solving; able to handle complex, ambiguous issues.
🎁 Benefits
- Opportunity to work with Kubernetes-native AI infrastructure at scale.
- Global team collaboration across regions.
- Professional development and learning opportunities.
- Open, collaborative work environment with cutting-edge tech.
- Competitive compensation and benefits package.
Back to all jobs