Senior LLM Inference Performance and Evaluation Engineer
You build benchmark pipelines for LLM inference performance, create model launch gates, maintain representative workloads, compare model and runtime options, automate regression detection, and collaborate with runtime engineers and SRE to identify bottlenecks and define service objectives and alert thresholds.
Responsibilities
- Build benchmark pipelines for latency, throughput, concurrency, error rate, and GPU utilization
- Create model launch gates for API compatibility, streaming, tool calling, reasoning, multimodal behavior, and long-context cases
- Maintain synthetic, replayed, and customer-like workloads
- Compare model, runtime, and provider options and recommend routing, fallback, pricing, and capacity decisions
- Automate regression detection in CI/CD and staging
- Identify bottlenecks and verify performance improvements
- Convert benchmark results into SLOs and alert thresholds
- Build dashboards, reports, and release gates
Requirements
- 5+ years of experience in ML infrastructure, performance engineering, model evaluation, QA automation, or backend testing for production systems
- Experience with TTFT, TPOT/ITL, request latency, token throughput, concurrency, and GPU-utilization metrics
- Strong Python skills
- Go experience preferred
- Familiarity with OpenAI and Anthropic APIs, vLLM, Dynamo, SGLang, Triton-style servers, and Kubernetes test environments
- Ability to design statistically meaningful tests and communicate tradeoffs
- Experience building dashboards, reports, and release gates