Staff / Principal Machine Learning Engineer, Serving
You will take models from the research team, containerize them, optimize their serving, and ensure they run reliably in production. You will work on inference optimization, model acceleration, and high-performance distributed systems to power realtime multimodal AI at scale for hundreds of millions of end users. You will question architectures, design benchmarks and prototypes to find answers, and treat performance, latency, and reliability as first-class product features rather than a box to check before launch.
Responsibilities
- Optimize realtime inference and serving frameworks
- Containerize models from the research team and ensure reliable production deployment
- Apply quantization, distillation, caching, continuous batching, paged attention, and speculative decoding
- Profile code and optimize performance on NVIDIA GPUs
- Handle multi-GPU/multi-node inference and thousands of concurrent connections
- Design benchmarks and prototypes to validate technical decisions
- Share work and contribute to open-source projects
Requirements
- Deep understanding of modern serving frameworks and techniques like vLLM or TRT-LLM
- Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
- Proficiency in C++, CUDA, Rust, or highly optimized Python
- Experience with Kubernetes, Ray, custom load balancing, multi-GPU/multi-node inference, and handling thousands of concurrent connections
- Non-trivial systems programming projects, open-source contributions to major inference engines, or deep-dive technical write-ups
- Full-cycle ownership from research model to production serving
- PhD in CS, Physics, Math, or equivalent practical experience building backend or ML systems
- Legal right to work in the United Kingdom
Benefits
- Equity