Senior Site Reliability Engineer (SRE, Compute Node Team)
JobgetherJob Description
📋 Description Ensure VM compute nodes are reliable, available, and fast. Analyze and troubleshoot Linux systems (user and kernel space). Resolve production issues with CPU, memory, NUMA, cgroups, and scheduling. Hands-on with virtualization tech including QEMU/KVM. Design observability at infra layer: metrics, logs, traces, alerts. Lead incident response, root-cause analysis, and post-incident improvements. 🎯 Requirements Deep Linux expertise: user space, kernel space, sched, memory, cgroups, and namespaces. Strong system architecture understanding across infrastructure layers. Hands-on virtualization with QEMU/KVM, including VM lifecycle management. Practical experience with containers, namespaces, and resource isolation. Strong debugging and hypothesis-driven incident investigation. Solid understanding of SRE principles: reliability engineering, system design, and ownership. 🎁 Benefits Competitive compensation package. Flexible working environment with ownership over projects and responsibilities. Opportunities for career growth and continuous learning. Chance to work on impactful AI infrastructure projects. Collaborative international environment with talented engineering teams. Opportunity to solve challenging technical problems at large scale.