Skip to content
ITA Jobs
All vacancies
Featherless AI logo

Machine Learning Engineer - Inference Optimization

Featherless AI

Remote · worldwide, UTC+4 acceptedSalary not disclosedfull-timeVerified recentlyHimalayas

We’re looking for a Machine Learning Engineer to own and push the limits of model inference performance at scale. You’ll work at the intersection of research and production—turning cutting-edge models into fast, reliable, and cost-efficient systems that serve real users.

Responsibilities

  • Optimize inference latency, throughput, and cost for large-scale ML models in production
  • Profile and bottleneck GPU/CPU inference pipelines (memory, kernels, batching, IO)
  • Implement and tune techniques such as
  • Quantization (fp16, bf16, int8, fp8)
  • KV-cache optimization & reuse
  • Speculative decoding, batching, and streaming
  • Model pruning or architectural simplifications for inference
  • Collaborate with research engineers to productionize new model architectures
  • Build and maintain inference-serving systems (e.g. Triton, custom runtimes, or bespoke stacks)
  • Benchmark performance across hardware (NVIDIA / AMD GPUs, CPUs) and cloud setups

Languages

    Work format
    Remote
    Seniority
    Mid
    Posted
    24 Sept 2026 (3 days ago)
    Last verified
    27 Sept 2026
    Apply by
    23 Nov 2026

    Keep looking