Skip to main content
Agents is in private beta, so things may change. If you hit a problem, tell us via Support & feedback.
An evaluation answers one question: did that change make my Agent better? It answers every question in your test cases, grades the replies with your scorers, and gives you a score you can compare.
An evaluation doesn’t run your Agent for real. It writes each reply from the version’s instructions and model alone, in one pass, with no tools, no integrations, and no memory of earlier messages. That makes it a good way to measure your instructions, and the wrong way to check whether the Agent uses its tools correctly. For that, read a real run in Run history.
Go to Agents → open an Agent → EvaluateEvaluation runsStart evaluation.
A completed evaluation run showing the per-scorer score and each case's ground truth beside the Agent's output

How do I run an evaluation?

1

Name it

Optional, but a name like Tightened refusal tone makes the list much easier to read later.
2

Pick the version to test

Whose instructions and model the replies are written from. The newest by default. See Versions.
3

Pick your scorers

Choose up to five scorers to grade the replies.
4

Choose the questions

Pick one or more datasets, then narrow by tag if you want. Hercules tells you how many cases it can run with the scorers you picked.
5

Start it

Click Run evaluation.

How do I prove a change helped?

Save a starting point, then compare against it.
1

Run an evaluation before you change anything

On the run’s page, click Set as baseline. That’s the score you’ll compare everything against.
2

Make your change and publish it

3

Run the same evaluation again

Open the new run and compare it against your baseline.
The comparison lines up each question against last time, so you can see exactly which answers got better and which got worse. The report lists every question, the answer you expected, the reply the evaluation produced, and each scorer’s grade. Filter it by All, Passed, Failed, or Errors.
Because you pick which version answers, you can compare two versions without publishing either one in between.

Why can’t some of my cases run?

Because one of your scorers compares the reply against the answer you expected, and those cases don’t have one. Cases without an expected answer get skipped. Add one on the Test cases page, or pick scorers that don’t need it.

What does an evaluation cost?

Two things cost credits: writing a reply to every question, and grading every reply. A big set of questions with several scorers isn’t cheap, so filter down by tag while you’re still tweaking things.

Additional FAQ

Yes, by status: In progress, Complete, Failed, or Canceled.
Its scores and results are gone. If it was your baseline, comparisons go back to having none. This can’t be undone.
Yes, 100. If you’ve picked more than that, Hercules tells you and you can’t start. Narrow it down by tag, or pick fewer datasets.
Both, for different jobs. An evaluation measures a change you already made. Auto-improve reads your run history and suggests changes to make.