Staff / Principal Machine Learning Engineer, Serving
You'll take models from the research team and own their full lifecycle, from containerizing them to optimizing their serving so they run reliably in production at scale. You'll work on sub-second multimodal inference systems, applying techniques like quantization, distillation, caching, continuous batching, paged attention, and speculative decoding to push performance further. You'll build and maintain distributed serving infrastructure, working with Kubernetes, Ray, and custom load balancing to handle multi-GPU and multi-node inference for thousands of concurrent connections. You'll profile code deeply and squeeze maximum performance out of NVIDIA GPUs using C++, CUDA, Rust, or highly optimized Python. You'll question architecture decisions when you see a better path to solving core latency or throughput problems, and you'll design benchmarks and prototypes to find answers when you don't know something yet. You'll share your work openly, contributing to open-source projects and technical write-ups that move the field forward.
Responsibilities
- Take models from the research team to production
- Containerize and optimize model serving
- Optimize inference using quantization, distillation, caching, continuous batching, paged attention, and speculative decoding
- Build distributed serving infrastructure across multiple GPUs and nodes
- Handle thousands of concurrent connections reliably
- Profile code and optimize performance on NVIDIA GPUs
- Design benchmarks and prototypes to resolve open questions
- Contribute to open-source projects and technical write-ups
Requirements
- Deep understanding of modern serving frameworks like vLLM or TRT-LLM
- Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
- Proficiency in C++, CUDA, Rust, or highly optimized Python
- Experience with Kubernetes, Ray, custom load balancing, and multi-GPU/multi-node inference
- Public work such as open-source contributions or technical write-ups
- Full-cycle ownership experience taking models to production
- PhD in CS, Physics, Math, or equivalent practical experience
- Professional fluency in English
- Legal right to work in Switzerland
Benefits
- Remote work
- Full U.S. visa and relocation support may be available for those interested in relocating to the San Francisco Bay Area