Site Reliability Engineer (SRE)
HelloKindredJob Description
📋 Description Operate and enhance Kubernetes platforms across AWS, Azure, and on-premise environments. Lead incident response, problem management, and root cause analysis activities. Manage cluster lifecycle: upgrades, patching, node pools, CNI/CSI, ingress, Rancher. Own observability strategy including dashboards, alerting, monitoring, and SLOs/SLIs. Implement GitOps practices using Fleet and reduce toil through automation. Apply secure API gateway and Web Application Firewall patterns. 🎯 Requirements Deep expertise in Kubernetes, Rancher, GitOps, Linux, and cloud networking. Strong experience operating in hybrid cloud environments across AWS, Azure, and on-premise. Strong automation and scripting skills in Python, Go, Bash, PowerShell, or .NET. Proven experience with Infrastructure as Code using Terraform and Crossplane. Experience implementing and managing observability tooling including Grafana, Prometheus, Jaeger or Tempo, CloudWatch, Loki, and OpenTelemetry. Experience operating within regulated environments including PCI DSS and GDPR.