Job Description
📋 Description Design, build, and maintain infra and tools for the ML lifecycle from development to deployment Operate production-grade Kubernetes cluster including CNI networking and storage Deploy ML orchestration tooling (e.g., Metaflow) and storage for reproducible workflows Validate GPU node health across multiple hardware types and standardize drivers/CUDA Port workflows, migrate datasets/artifacts, and roll out tool updates with engineering teams Write runbooks, architecture docs, and disaster recovery procedures 🎯 Requirements 3+ years in infrastructure, DevOps, or MLOps roles or equivalent Production Kubernetes experience (networking, storage, RBAC, troubleshooting) Python and/or Bash scripting; IaC tools (Ansible, Helm, Terraform, or similar) Hands-on with distributed storage (S3, MinIO, NetApp, or similar) CI/CD pipelines experience (GitLab CI, ArgoCD, or equivalent) Linux administration and networking fundamentals; strong communication/doc skills 🎁 Benefits Onsite weekly schedule with set hours; close collaboration with cross-functional teams Equal opportunity employer; vaccination context noted; covered by company policies