Staff / Principal Machine Learning Engineer, Serving

You will take models from the research team, containerize them, optimize their serving, and ensure they run reliably in production. You will work on inference optimization, model acceleration, and high-performance distributed systems to power realtime multimodal AI at scale for hundreds of millions of end users. You will question architectures, design benchmarks and prototypes to find answers, and treat performance, latency, and reliability as first-class product features rather than a box to check before launch.

Responsibilities

  • Optimize realtime inference and serving frameworks
  • Containerize models from the research team and ensure reliable production deployment
  • Apply quantization, distillation, caching, continuous batching, paged attention, and speculative decoding
  • Profile code and optimize performance on NVIDIA GPUs
  • Handle multi-GPU/multi-node inference and thousands of concurrent connections
  • Design benchmarks and prototypes to validate technical decisions
  • Share work and contribute to open-source projects

Requirements

  • Deep understanding of modern serving frameworks and techniques like vLLM or TRT-LLM
  • Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
  • Proficiency in C++, CUDA, Rust, or highly optimized Python
  • Experience with Kubernetes, Ray, custom load balancing, multi-GPU/multi-node inference, and handling thousands of concurrent connections
  • Non-trivial systems programming projects, open-source contributions to major inference engines, or deep-dive technical write-ups
  • Full-cycle ownership from research model to production serving
  • PhD in CS, Physics, Math, or equivalent practical experience building backend or ML systems
  • Legal right to work in the United Kingdom

Benefits

  • Equity

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available