Manager, AI Operations
JobgetherNorth AmericaFull TimeEngineering
Remotely
securityincident responsemlopsproductiongovernanceai observability
Job Description
📋 Description
- Oversee deployment support, day-to-day reliability, monitoring, and production ops of AI systems
- Lead incident response, troubleshooting, root-cause analysis, and remediation for AI/ML services
- Partner with MLOps, Cloud Engineering, Infrastructure, Security, and IT to establish robust
- Build and maintain AI observability: drift detection, latency, uptime, logging, alerting, and
- Define production health metrics and dashboards for real-time visibility into performance and risks.
- Coordinate transitions from development to production with production-readiness reviews and
🎯 Requirements
- 7+ years in AI/ML ops, MLOps, production data science, or related; at least 2 years of people
- Hands-on experience operating traditional ML and Generative AI in production: deployment
- Strong observability practices: model drift, performance monitoring, logging, alerting, uptime
- Experience partnering with Security and Infrastructure on production risk, access controls, and
- Excellent communication and stakeholder management; ability to translate risks into business terms.
- Proven ability to lead teams, manage capacity, and establish scalable operational processes.
🎁 Benefits
- Comprehensive medical, dental, and vision insurance.
- 401(k) with company matching.
- Flexible paid time off and paid parental leave.
- HSA and FSA options; educational reimbursement.
- Employer-assisted mental health support and professional development opportunities.
- Collaborative culture with opportunities to impact AI ops and healthcare efficiency.