ML Operations Engineer
DeepIntentBelgradeFull TimeEngineering
Remotely
pythondockerkubernetessparkairflowmlflowkubeflowargo
Job Description
📋 Description
- Partner with Data Science and AI Engineering teams to adopt MLOps best practices and migrate
- Implement and maintain model tracking, versioning and experiment management (MLflow) with
- Build CI/CD pipelines purpose-built for ML/AI artifacts (model registries, container image
- Continuously improve platform reliability, cost-efficiency and maintainability of the underlying
- Establish monitoring and observability for ML/AI systems (Prometheus, Grafana) covering GPU
- Design and operate ML/AI deployment infrastructure, including GPU cluster architecture, model
🎯 Requirements
- Bachelor's degree in Computer Science or similar technical field of study, or equivalent practical
- Strong software engineering skills in complex, distributed, multi-language systems (Python
- Hands-on experience with Spark, Docker and Kubernetes in production environments
- Experience building and operating end-to-end distributed systems
- Experience developing and maintaining ML systems built with open-source MLOps tools (e.g., MLflow
- Strong understanding of software testing, benchmarking, and CI/CD practices
🎁 Benefits
- Competitive base salary plus performance-based bonus
- Comprehensive medical insurance
- Flexible PTO
- Hybrid-friendly culture with flexible work options
- Professional development reimbursement
- WiFi reimbursement
Back to all jobs