Senior LLM Inference Performance and Evaluation Engineer

You build benchmark pipelines for LLM inference performance, create model launch gates, maintain representative workloads, compare model and runtime options, automate regression detection, and collaborate with runtime engineers and SRE to identify bottlenecks and define service objectives and alert thresholds.

Responsibilities

  • Build benchmark pipelines for latency, throughput, concurrency, error rate, and GPU utilization
  • Create model launch gates for API compatibility, streaming, tool calling, reasoning, multimodal behavior, and long-context cases
  • Maintain synthetic, replayed, and customer-like workloads
  • Compare model, runtime, and provider options and recommend routing, fallback, pricing, and capacity decisions
  • Automate regression detection in CI/CD and staging
  • Identify bottlenecks and verify performance improvements
  • Convert benchmark results into SLOs and alert thresholds
  • Build dashboards, reports, and release gates

Requirements

  • 5+ years of experience in ML infrastructure, performance engineering, model evaluation, QA automation, or backend testing for production systems
  • Experience with TTFT, TPOT/ITL, request latency, token throughput, concurrency, and GPU-utilization metrics
  • Strong Python skills
  • Go experience preferred
  • Familiarity with OpenAI and Anthropic APIs, vLLM, Dynamo, SGLang, Triton-style servers, and Kubernetes test environments
  • Ability to design statistically meaningful tests and communicate tradeoffs
  • Experience building dashboards, reports, and release gates

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available