All vacancies
GPU Cluster Architect
Nebius
Remote · United StatesSalary not disclosedfull-timeVerified recentlyOver a month oldNebius Careers
We are seeking a GPU Cluster Architect to drive the design of our next-generation AI infrastructure. In this high-impact, hands-on role, you will make end-to-end architectural decisions across compute, networking, and storage — ensuring our platforms can meet the massive scale, performance, and reliability requirements of modern AI workloads.
Responsibilities
- Design: Architect scalable GPU cluster topologies including compute nodes, interconnect (InfiniBand, Ethernet), storage, and control planes.
- Performance Modeling
- Analyze AI/ML workloads (e.g. LLM training, inference) to inform design tradeoffs across latency, bandwidth, and GPU density.
- Network Architecture: Align with network architect relevant design and validate low-latency, high-throughput interconnects (e.g., InfiniBand HDR/NDR, RoCEv2) at POD and DC scale.
- Storage Integration: Work with storage teams to optimize performance for training datasets, checkpointing, and others.
- Reliability & Monitoring: Understand and analyze signal from monitoring systems to the detect flows in design
- Collaboration
- Partner with site reliability, networking, storage, and DC engineering teams to operationalize and scale your architecture.
Languages
- Work format
- Remote
- Seniority
- Senior
- Posted
- 21 Aug 2025 (13mo ago)
- Last verified
- 4 Oct 2026
