Santa ClaraFull TimeEngineering
Remotely
awsai/mlkubernetesazureterraformobservabilitygitopsauto remediation
Job Description
📋 Description
- Design, deploy, and operate enterprise-scale Kubernetes clusters across hybrid and multi-cloud
- Architect closed-loop auto-remediation to detect and resolve infra failures with minimal human
- Develop SRE tooling for monitoring, incident management, logs, and observability for global on-call
- Establish SLOs, error budgets, and runbooks to balance rapid response with automation-driven
- Implement Infrastructure-as-Code and GitOps pipelines for reproducible deployments across
- Lead on-call schedules across time zones and mentor engineers in reliability best practices.
Back to all jobs