Senior / Lead Machine Learning Engineer, Serving

You will work on optimizing realtime inference and serving of state-of-the-art voice models at massive scale. You'll take models from the research team, containerize them, optimize their serving, and ensure they run reliably in production. You'll dive deep into inference optimization, model acceleration, and distributed systems, squeezing every ounce of performance out of NVIDIA GPUs while handling thousands of concurrent connections. You'll be trusted to pick a direction and build the map as you go, treating performance, latency, and reliability as first-class product features. You'll collaborate daily with US-based leadership and engineering teams, question architectures when you see a better way to solve latency or throughput problems, and ship code that is stable and visible.

Responsibilities

  • Optimize realtime inference and model serving at scale
  • Take models from the research team, containerize them, and optimize their serving
  • Ensure models run reliably in production
  • Apply model acceleration techniques such as quantization, distillation, caching, continuous batching, paged attention, and speculative decoding
  • Profile code and optimize performance on NVIDIA GPUs
  • Handle distributed systems and scaling using Kubernetes, Ray, and custom load balancing
  • Manage multi-GPU and multi-node inference reliably handling thousands of concurrent connections
  • Collaborate daily with US-based leadership and engineering teams

Requirements

  • Deep understanding of modern serving frameworks and techniques like vLLM or TRT-LLM
  • Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
  • Proficiency in C++, CUDA, Rust, or highly optimized Python
  • Experience with Kubernetes, Ray, custom load balancing, multi-GPU/multi-node inference
  • Experience reliably handling thousands of concurrent connections
  • Non-trivial systems programming projects, open-source contributions to major inference engines, or deep-dive technical write-ups
  • Ability to take a model from research to production, containerizing and optimizing its serving
  • PhD in CS, Physics, Math, or equivalent practical experience building backend or ML systems
  • Professional fluency in English, written and spoken

Benefits

  • Full U.S. visa and relocation support may be available for candidates interested in relocating to the San Francisco Bay Area

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available