Senior ML Ops Engineer
KAYAKRemotely
pythondockerkubernetesmlopsobservabilitymodel servingprometheusgrafana
Job Description
📋 Description
- Build and maintain ML infrastructure end-to-end: extend and operate CI/CD pipelines, model
- Own model deployment and serving: define and evolve tooling for model serving with low latency and
- Develop core MLOps capabilities: feature stores, model registries, automated monitoring for
- Operationalize infrastructure for the ML team: enable Kubernetes autoscaling and GPU provisioning
- Improve platform reliability and performance: design resilient monitoring and automate for
- Empower Data Scientists through standardized workflows: build golden paths to streamline model
🎯 Requirements
- Experience building and operating ML platforms in production environments.
- Strong knowledge of containerization and orchestration (Docker, Kubernetes), Linux internals, and
- Familiarity with ML lifecycle tooling including orchestration frameworks, feature stores, model
- Experience owning production systems: define SLOs, build observability, participate in incident
- Comfort coding production-quality Python or similar language.
- Experience modernizing production infrastructure with a focus on reliability, risk, and cost.
🎁 Benefits
- Work from (almost) anywhere for up to 20 days per year.
- Mental health and well-being perks: therapy, HeadSpace, company-wide time off, no meetings on
- Paid parental leave and paid volunteer time.
- Career growth: Development Dollars, leadership development, on-demand e-learnings.
- Travel discounts, ERGs, and social/team events.
- Office in Friedrichshain, Berlin.