All vacancies
ML Infrastructure Engineer
Nebius
Remote · Europe · United StatesSalary not disclosedfull-timeVerified recentlyOver a month oldNebius Careers
We are seeking a highly skilled ML/AI Engineer to join our team to lead and support benchmarking of GPU platforms benchmarking of GPU platforms for machine learning and AI workloads. You will play a critical role in evaluating the performance of GPU-based hardware for various deep learning and AI frameworks, enabling data-driven decisions for platform optimisation and next-generation hardware development.
Responsibilities
- Work closely with hardware, development teams to profile and analyse GPU performance at the system and kernel level.
- Evaluate and compare GPU performance across different platforms, architectures, and software stacks (e.g.,CUDA, ROCm).
- Debug and optimise ML workloads to run efficiently on GPU hardware, identifying and resolving performance bottlenecks.
- Perform acceptance testing acceptance testing for new GPU clusters, ensuring hardware and software meet performance, stability, and compatibility requirements for AI workloads.
- Perform experiments across diverse GPU system configurations to assess the impact of varying interconnect strategies and system-level optimisations on performance and scalability.
- Develop tools and dashboards to visualise performance metrics visualise performance metrics, bottlenecks, and trends.
- Contribute to internal tooling, frameworks, and best practices
Languages
- Work format
- Remote
- Seniority
- Mid
- Posted
- 12 May 2026 (4mo ago)
- Last verified
- 4 Oct 2026
