All vacancies
Senior Machine Learning Engineer, LLM Inference Optimization
Nebius
Remote · worldwide not confirmedSalary not disclosedfull-timeVerified recentlyHimalayas
Nebius Token Factory is building fast, reliable, and cost-efficient inference services for frontier models. As a Senior Machine Learning Engineer on our Applied AI team, you will own model and endpoint optimization from model artifacts through production deployment.
Responsibilities
- Own optimization work for specific model families, customer endpoints, or serving backends.
- Run engine comparisons and recommend practical serving configurations for specific workloads.
- Debug model quality or performance regressions during production rollouts.
- Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token.
- Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems.
- Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.
- Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.
- Build reproducible benchmark harnesses for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token.
- Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers.
- Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations.
Languages
- Work format
- Remote
- Seniority
- Senior
- Posted
- 27 Sept 2026 (today)
- Last verified
- 27 Sept 2026
- Apply by
- 26 Nov 2026
