Skip to content
ITA Jobs
All vacancies
Nebius logo

Senior Site Reliability Engineer (SRE, Compute Node Team)

Nebius

Remote · EuropeSalary not disclosedfull-timeVerified recentlyOver a month oldNebius Careers

We are looking for a Senior Site Reliability Engineer (SRE) to join the Compute Node team at Nebius AI Cloud. The Compute Node team is responsible for building and operating the cluster scheduler and node-level services that run and manage virtual machines across all cloud regions.

Responsibilities

  • Ensure reliability, availability and performance of compute nodes running VMs
  • Analyze and debug Linux systems across user space and kernel space, understanding capabilities, limitations and trade-offs at each layer
  • Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups and scheduling
  • Work hands-on with virtualization and containerization, primarily using QEMU/KVM and Linux-native technologies
  • Design and evolve observability as a core capability of the node layer: metrics, logs, traces, alerts, SLIs and SLOs
  • Lead incident response, root-cause analysis, and postmortems, driving long-term reliability improvements
  • Collaborate closely with platform, kernel/hypervisor, GPU and infrastructure teams to improve system design and operability

Languages

    Work format
    Remote
    Seniority
    Senior
    Posted
    26 Jan 2026 (8mo ago)
    Last verified
    4 Oct 2026

    Keep looking