Data Scientist, Agent
You will own the measurement and improvement of an AI agent. You will define quality metrics, build evaluation and experimentation systems, analyze agent traces and telemetry, identify regressions, and work with engineering to improve success rates, task completion, and error rates.
Responsibilities
- Define and own agent quality metrics
- Build evaluation systems and experiment frameworks
- Determine whether agent changes should ship
- Analyze agent traces and telemetry
- Identify concrete fixes with the agent engineering team
- Build tooling and agents for continuous evaluations
- Set standards for judging agent behavior without an answer key
Requirements
- Experience or strong interest in LLM evaluation and observability
- Strong SQL and Python
- Applied statistics
- Experimentation and A/B testing
- Ability to design experiments with noisy outcomes
- Ability to measure agent behavior without a clean answer key
- Entrepreneurial approach
- Ability to work closely with agent engineers