Job Description
📋 Description Build and operate reliable infrastructure for research workloads and research-facing services. Improve data infra stack: processing, crawl/ingest, caching, search, observability, and cluster services. Improve cluster bootstrapping, provisioning, automation, and deployment workflows. Debug issues across networking, compute, storage, orchestration, and reliability layers. Build software and automation to reduce manual work and improve reliability. Partner with researchers and engineers to translate needs into durable solutions. 🎯 Requirements Strong systems fundamentals and scalable infrastructure understanding. Proficient with Linux, networking, Kubernetes, provisioning, and distributed systems. Able to write software to automate, debug, and improve infrastructure. Strong execution mindset; can drive ambiguous infra work independently. Enjoy supporting a wide surface area of systems, from tooling to platform services. Pragmatic about building vs using existing tools. 🎁 Benefits PXE boot, cluster provisioning, bare-metal infra, large-scale fleet management. Operating Kubernetes or similar orchestration systems at scale. Infrastructure-as-code, CI/CD, observability, deployment automation. Supporting search infra, data platforms, ingest systems, or large-scale research workflows. Git-based workflows and internal developer tooling. Reliability, scale, and speed-focused environments.