Job Description
📋 Description Architect scalable, self-healing infra across multi-region deployments using Kubernetes. Drive AI enablement across engineering with Claude Code, Cursor, and Codex. Build AI tooling and automation (K8s upgrades, IaC analysis, n8n). Build and maintain internal developer platform (IDP) for self-service and reliability. Develop observability using Prometheus and Grafana for metrics, dashboards, alerts. Lead incident management with blameless postmortems; enforce SLIs, SLOs, budgets. 🎯 Requirements 5+ years as Infrastructure Engineer focused on reliability (SRE/Platform). Experience driving company-wide reliability efforts with SLO frameworks. Strong proficiency with OpenTelemetry, Prometheus, Grafana. Deep experience with AWS/GCP, Kubernetes, and multi-region architectures. Skilled with Terraform, Helm, and GitOps (ArgoCD) automation-first. Experience leveraging agentic development tools (Claude Code, Cursor, Codex) and n8n. Solid networking fundamentals — VPC, DNS, IPAM, security groups, Istio. Strong cross-functional communicator across SRE, security, and product. Blockchain infra, distributed systems, or high-throughput RPC experience a plus. 🎁 Benefits Medical, Dental, & Vision Gym Reimbursement Home Office Build-out Budget In-Office Group Meals Wellbeing & Mental Health Perks Learning & Development Stipend