Staff / Principal Machine Learning Engineer, Serving

You'll take models from the research team and own their full lifecycle, from containerizing them to optimizing their serving so they run reliably in production at scale. You'll work on sub-second multimodal inference systems, applying techniques like quantization, distillation, caching, continuous batching, paged attention, and speculative decoding to push performance further. You'll build and maintain distributed serving infrastructure, working with Kubernetes, Ray, and custom load balancing to handle multi-GPU and multi-node inference for thousands of concurrent connections. You'll profile code deeply and squeeze maximum performance out of NVIDIA GPUs using C++, CUDA, Rust, or highly optimized Python. You'll question architecture decisions when you see a better path to solving core latency or throughput problems, and you'll design benchmarks and prototypes to find answers when you don't know something yet. You'll share your work openly, contributing to open-source projects and technical write-ups that move the field forward.

Responsibilities

  • Take models from the research team to production
  • Containerize and optimize model serving
  • Optimize inference using quantization, distillation, caching, continuous batching, paged attention, and speculative decoding
  • Build distributed serving infrastructure across multiple GPUs and nodes
  • Handle thousands of concurrent connections reliably
  • Profile code and optimize performance on NVIDIA GPUs
  • Design benchmarks and prototypes to resolve open questions
  • Contribute to open-source projects and technical write-ups

Requirements

  • Deep understanding of modern serving frameworks like vLLM or TRT-LLM
  • Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
  • Proficiency in C++, CUDA, Rust, or highly optimized Python
  • Experience with Kubernetes, Ray, custom load balancing, and multi-GPU/multi-node inference
  • Public work such as open-source contributions or technical write-ups
  • Full-cycle ownership experience taking models to production
  • PhD in CS, Physics, Math, or equivalent practical experience
  • Professional fluency in English
  • Legal right to work in Switzerland

Benefits

  • Remote work
  • Full U.S. visa and relocation support may be available for those interested in relocating to the San Francisco Bay Area

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available