All vacancies
Senior Site Reliability Engineer (SRE, Compute Node Team)
Nebius
Remote · EuropeSalary not disclosedfull-timeVerified recentlyOver a month oldNebius Careers
We are looking for a Senior Site Reliability Engineer (SRE) to join the Compute Node team at Nebius AI Cloud. The Compute Node team is responsible for building and operating the cluster scheduler and node-level services that run and manage virtual machines across all cloud regions.
Responsibilities
- Ensure reliability, availability and performance of compute nodes running VMs
- Analyze and debug Linux systems across user space and kernel space, understanding capabilities, limitations and trade-offs at each layer
- Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups and scheduling
- Work hands-on with virtualization and containerization, primarily using QEMU/KVM and Linux-native technologies
- Design and evolve observability as a core capability of the node layer: metrics, logs, traces, alerts, SLIs and SLOs
- Lead incident response, root-cause analysis, and postmortems, driving long-term reliability improvements
- Collaborate closely with platform, kernel/hypervisor, GPU and infrastructure teams to improve system design and operability
Languages
- Work format
- Remote
- Seniority
- Senior
- Posted
- 26 Jan 2026 (8mo ago)
- Last verified
- 4 Oct 2026
