Staff / Principal Machine Learning Engineer, Serving
You'll work on optimizing realtime, sub-second multimodal inference at scale for some of the world's top-ranked realtime voice models. You'll take models from the research team, containerize them, optimize their serving, and ensure they run reliably in production. You'll apply techniques like quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding to accelerate models, and you'll build high-performance systems in C++, CUDA, Rust, or highly optimized Python. You'll design and operate distributed systems using Kubernetes, Ray, custom load balancing, and multi-GPU/multi-node inference to reliably handle thousands of concurrent connections. You'll be expected to question architectures, prototype benchmarks to answer open questions, and treat performance, latency, and reliability as first-class product features.
Responsibilities
- Optimize realtime inference serving using frameworks like vLLM or TRT-LLM
- Apply model acceleration techniques such as quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
- Build high-performance systems in C++, CUDA, Rust, or optimized Python and profile code to maximize GPU performance
- Design and scale distributed systems with Kubernetes, Ray, custom load balancing, and multi-GPU/multi-node inference
- Handle thousands of concurrent connections reliably
- Take models from the research team, containerize them, optimize their serving, and ensure reliable production operation
- Contribute to open-source projects and share technical write-ups
Requirements
- Deep understanding of modern serving frameworks and techniques like vLLM or TRT-LLM
- Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
- Proficiency in C++, CUDA, Rust, or highly optimized Python
- Experience profiling code and optimizing performance on NVIDIA GPUs
- Experience with Kubernetes, Ray, custom load balancing, and multi-GPU/multi-node inference
- Experience reliably handling thousands of concurrent connections
- Non-trivial systems programming projects, open-source contributions to major inference engines, or deep-dive technical write-ups
- Ability to take a model from research to production, including containerization and optimization
- PhD in CS, Physics, Math, or equivalent practical experience building backend or ML systems
Benefits
- Relocation assistance
- Bonus
- Equity