Senior Site Reliability Engineer (SRE, Compute Node Team)
JobgetherJob Description
📋 Description Ensure reliability, availability, and performance of compute nodes across clouds. Analyze and troubleshoot Linux systems in user and kernel space. Investigate and resolve production issues involving CPU, memory, NUMA, and scheduling. Hands-on with virtualization tech, including QEMU/KVM. Design observability at infra layer: metrics, logs, traces, alerts, SLIs, SLOs. Lead incident response, perform root-cause analysis, and drive improvements. 🎯 Requirements Linux expertise across user/kernel space and subsystems (scheduling, memory, cgroups, namespaces). Strong system architecture knowledge and performance trade-offs across layers. Hands-on with virtualization (QEMU/KVM), VM lifecycle, and perf optimization. Practical experience with containers, namespaces, and resource isolation. Strong debugging with hypothesis-driven incident investigation. Solid grasp of SRE principles: reliability, design, and operational ownership. 🎁 Benefits Competitive compensation package. Flexible working environment with ownership over projects. Opportunities for career growth and continuous learning. Chance to work on impactful AI infrastructure projects. Collaborative international environment with talented engineering teams. Opportunity to solve challenging technical problems at large scale.