Senior / Lead Machine Learning Engineer, Serving

You'll work on optimizing realtime inference for state-of-the-art voice models, taking models from the research team and containerizing, optimizing, and ensuring they run reliably in production. You'll dive deep into inference optimization using frameworks like vLLM or TRT-LLM, apply model acceleration techniques such as quantization, distillation, caching, continuous batching, paged attention, and speculative decoding, and squeeze maximum performance out of NVIDIA GPUs. You'll build and scale distributed systems using Kubernetes, Ray, and custom load balancing to reliably handle thousands of concurrent connections across multi-GPU and multi-node setups. You'll be expected to question architectures, design benchmarks and prototypes to find answers, and treat performance, latency, and reliability as first-class product features rather than boxes to check before launch.

Responsibilities

  • Optimize realtime inference using serving frameworks like vLLM or TRT-LLM
  • Apply model acceleration techniques such as quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
  • Profile code and optimize performance on NVIDIA GPUs using C++, CUDA, Rust, or optimized Python
  • Build and scale distributed systems using Kubernetes, Ray, and custom load balancing
  • Handle multi-GPU and multi-node inference reliably for thousands of concurrent connections
  • Take models from the research team, containerize them, optimize their serving, and ensure reliable production operation
  • Design benchmarks and prototypes to validate architectural decisions
  • Collaborate daily with US-based leadership and engineering teams

Requirements

  • Deep understanding of modern serving frameworks and techniques like vLLM or TRT-LLM
  • Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
  • Proficiency in C++, CUDA, Rust, or highly optimized Python
  • Experience with Kubernetes, Ray, custom load balancing, multi-GPU/multi-node inference
  • Non-trivial systems programming projects, open-source contributions to major inference engines, or deep-dive technical write-ups
  • Full-cycle ownership experience taking a model from research to production
  • PhD in CS, Physics, Math, or equivalent practical experience building backend or ML systems
  • Professional fluency in English (written and spoken)

Benefits

  • Full U.S. visa and relocation support may be available for candidates interested in relocating to the San Francisco Bay Area

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available