ML Ops Engineer
AnaplanRemotely
dockerkubernetesansibleterraformprometheusgrafanamlflowray
Job Description
📋 Description
- Design, scale, and maintain high-performance MLOps and LLMOps infrastructure for an AI-infused
- Collaborate with Data Scientists, ML Engineers, and Cloud Infra teams to streamline training
- Ensure GPU utilisation, reliability, and cost-efficiency across cloud environments
🎯 Requirements
- Hands-on production experience in DevOps, SRE, or Platform Engineering for AI/ML infra
- Proven track record deploying and operationalising ML models and LLMs in cloud-native production
- Experience managing GPU infrastructure and HPC environments
- Advanced proficiency in Kubernetes, Docker, Helm, KubeFlow; service meshes (e.g., Istio)
- Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins
- Experience with vLLM, Ray, MLflow, LangChain / LangSmith, DeepSpeed, or Hugging Face TGI
🎁 Benefits
- DEIB focused culture and opportunities to grow within a global team
- Competitive benefits and career development
- Work on cutting-edge AI/ML infrastructure and platform engineering
Back to all jobs