Infrastructure Site Reliability Engineer
RadiantJob Description
📋 Description Deploy resilient, scalable infra for AI/HPC workloads Optimize Linux config, BIOS/firmware, kernel, and disk I/O Manage bare-metal infra with IPMI, Redfish, etc Build automation and IaC for platform lifecycle Provide 24x7 prod support with on-call rotations Mentor junior engineers and share knowledge 🎯 Requirements 5+ years in globally scaled, 24/7 production environments Expert Linux administration (Ubuntu distributions) System tuning, disk I/O optimization, hardware perf tweaks Familiar with IPMI/Redfish/PXE for out-of-band management Automation: Bash, Python, Ansible Observability: Prometheus, Grafana; Orchestration: Kubernetes 🎁 Benefits 25 days annual leave Learning Time to focus on new skills Private medical insurance via Bupa Cycle to Work Scheme Gympass subscription Participation in the company shares program