Senior / Lead Machine Learning Engineer, Serving
You will work on optimizing realtime inference and serving of state-of-the-art voice models at massive scale. You'll take models from the research team, containerize them, optimize their serving, and ensure they run reliably in production. You'll dive deep into inference optimization, model acceleration, and distributed systems, squeezing every ounce of performance out of NVIDIA GPUs while handling thousands of concurrent connections. You'll be trusted to pick a direction and build the map as you go, treating performance, latency, and reliability as first-class product features. You'll collaborate daily with US-based leadership and engineering teams, question architectures when you see a better way to solve latency or throughput problems, and ship code that is stable and visible.
Responsibilities
- Optimize realtime inference and model serving at scale
- Take models from the research team, containerize them, and optimize their serving
- Ensure models run reliably in production
- Apply model acceleration techniques such as quantization, distillation, caching, continuous batching, paged attention, and speculative decoding
- Profile code and optimize performance on NVIDIA GPUs
- Handle distributed systems and scaling using Kubernetes, Ray, and custom load balancing
- Manage multi-GPU and multi-node inference reliably handling thousands of concurrent connections
- Collaborate daily with US-based leadership and engineering teams
Requirements
- Deep understanding of modern serving frameworks and techniques like vLLM or TRT-LLM
- Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
- Proficiency in C++, CUDA, Rust, or highly optimized Python
- Experience with Kubernetes, Ray, custom load balancing, multi-GPU/multi-node inference
- Experience reliably handling thousands of concurrent connections
- Non-trivial systems programming projects, open-source contributions to major inference engines, or deep-dive technical write-ups
- Ability to take a model from research to production, containerizing and optimizing its serving
- PhD in CS, Physics, Math, or equivalent practical experience building backend or ML systems
- Professional fluency in English, written and spoken
Benefits
- Full U.S. visa and relocation support may be available for candidates interested in relocating to the San Francisco Bay Area