Staff AI Engineer Model Post-Training and Alignment

You will design, execute, and optimize post-training pipelines for large language models across data strategy, reward modeling, reinforcement learning, alignment, evaluation, and production inference. You will develop domain-specific data pipelines, train specialized models, build RLAIF systems, optimize serving, and productionize training and deployment workflows.

Responsibilities

  • Lead and execute LLM post-training pipelines
  • Design DPO and GRPO training paradigms
  • Develop domain-specific data recipes and augmentation pipelines
  • Train specialized small models from scratch
  • Build and refine reward models
  • Design and implement RLAIF closed-loop systems
  • Optimize inference efficiency and deploy models
  • Evaluate model performance with benchmarks and feedback loops
  • Collaborate to productionize training and deployment workflows

Requirements

  • Bachelor's degree in Computer Science, AI, Machine Learning, or a related field
  • 8+ years of industry experience
  • Large model post-training
  • Preference learning
  • DPO
  • GRPO
  • Reinforcement learning
  • Domain-specific data strategies
  • Training specialized small models from scratch
  • Reward modeling
  • RLAIF
  • vLLM
  • SGLang
  • Low-latency production deployment

Benefits

  • L&D programs
  • Education subsidy
  • Team building programs
  • Company events
  • Wellness allowances
  • Meal allowances
  • Comprehensive healthcare schemes for employees and dependants

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available