All vacancies
Machine Learning Engineer - Inference Optimization
Featherless AI
Remote · worldwide, UTC+4 acceptedSalary not disclosedfull-timeVerified recentlyHimalayas
We’re looking for a Machine Learning Engineer to own and push the limits of model inference performance at scale. You’ll work at the intersection of research and production—turning cutting-edge models into fast, reliable, and cost-efficient systems that serve real users.
Responsibilities
- Optimize inference latency, throughput, and cost for large-scale ML models in production
- Profile and bottleneck GPU/CPU inference pipelines (memory, kernels, batching, IO)
- Implement and tune techniques such as
- Quantization (fp16, bf16, int8, fp8)
- KV-cache optimization & reuse
- Speculative decoding, batching, and streaming
- Model pruning or architectural simplifications for inference
- Collaborate with research engineers to productionize new model architectures
- Build and maintain inference-serving systems (e.g. Triton, custom runtimes, or bespoke stacks)
- Benchmark performance across hardware (NVIDIA / AMD GPUs, CPUs) and cloud setups
Languages
- Work format
- Remote
- Seniority
- Mid
- Posted
- 24 Sept 2026 (3 days ago)
- Last verified
- 27 Sept 2026
- Apply by
- 23 Nov 2026
