ova

Evaluation

Evals for agentic workflows: build them into software you own

Evals for agentic workflows belong inside the product: traces, human approval steps and a regression set you own decide how much autonomy each step earns.

By Frideric Pétré · · 8 min read

A workflow line with an approval step in the middle; corrections from that step drop into a stack of checked cases, the regression set, which loops back to the start of the workflow.

Evals for agentic workflows work best when they are built into the product, not run once beside it. The workflow records what each step did, people approve or correct the result at defined points, and those corrections accumulate into a regression set that is rerun before every change. That set then decides how much autonomy each step is given, and it belongs to the company that owns the software.

This article looks at the building side: what has to exist in the software for evaluation to keep working after launch.

What are evals for agentic workflows?

Evals for agentic workflows are repeatable checks that run recorded or representative cases through a workflow and compare what each step produced, and which actions it took, with what the business has accepted as correct.

Three things separate them from a general model benchmark. They test your workflow, with your data, tools and rules, and not a model in isolation. They look at steps and actions as well as the final text. And they are rerun whenever something changes, which is why they need to live close to the code.

Anthropic's engineering team draws a useful line between capability evals, which ask what an agent can do well, and regression evals, which ask whether it still handles the tasks it used to (Demystifying evals for AI agents). For a business workflow in production, the second question is the one that protects daily operations.

Why is a prototype check not a production eval?

A prototype is usually judged by a few people trying a few cases. That is a reasonable way to decide whether an idea deserves a build. It says little about how the workflow behaves months later, after a model update, a new data source and several prompt changes. A prototype is only the beginning, and evaluation is one of the places where the distance to an operational product shows.

Prototype check Production eval
Cases Chosen by the builder, often the ones that work Drawn from real runs, including exceptions and past errors
Expected result In someone's head Written down and accepted by the process owner
When it runs Once, before the demo Before every change, and on a schedule
What it inspects The final answer Each step: context used, actions taken, result
Who can repeat it The person who built it Anyone with access to the repository
What a failure triggers A prompt tweak A blocked release, or a step moved back to review

Nothing in the right-hand column is exotic. It needs three things the software must provide: a record of each run, a place where people correct the work, and a way to keep those corrections.

Agent evaluation in production starts with the audit trail

You cannot evaluate what you did not record. A useful trace for a workflow step holds the input, the context that was retrieved and its source, the tools called and with which arguments, the proposed result, and the versions of the prompt, model and configuration in use. With that, a run from last Tuesday can be replayed against next month's version.

This is the same record that makes actions and errors traceable for the business. An audit trail for AI agents and the raw material for evals are one data set seen from two angles. European regulation points the same way for the systems it classes as high-risk: the EU AI Act requires that they technically allow for the automatic recording of events over their lifetime, and that they can be effectively overseen by people while in use. Whether or not a given workflow falls in that class, logging and oversight are sound engineering.

The core of Nova Backbone provides workflow orchestration with approval steps and audit trails. To be precise about what that means: the Backbone is not an evaluation product. It produces the traces and corrections that evals are built from. Which cases are kept, how they are checked and when they run is agreed per project and built as part of your product.

Human approval steps are where the standard gets written down

Process documentation rarely captures everything experienced colleagues know. Their corrections do. Each time a reviewer approves, edits or rejects a proposed step, they state what correct looks like for that case.

The design of the approval step determines whether that knowledge is kept or lost:

  • Approve records that this output, for this input, was acceptable.
  • Edit records the difference between what was proposed and what was actually used.
  • Reject records a reason, ideally chosen from a short list: wrong source, missing information, outside policy, should have escalated.

The reviewer does their normal job and the software keeps the evidence. This is also why an assistant such as Nova Companion prepares next steps that people confirm: the confirmation is a control and a source of examples at once. The same pattern appears in the illustrative custom applications on this site, where a human check or release is part of the workflow.

How does a correction become a regression set for AI agents?

The loop that turns production into evaluation: a production run leaves a trace, a reviewer corrects it, the correction joins the regression set, and the set is rerun before every change.

A regression set is a versioned collection of cases with known acceptable results. Each change to the workflow is run against it before release. Building one from production is routine work:

  1. Select. Take corrected and rejected runs, plus a sample of approved ones, so the set covers routine work as well as exceptions.
  2. Clean. Remove or mask personal and confidential data the case does not need, and apply the same access rules as the source systems.
  3. Fix the expected result. Store what the reviewer accepted, or the rule the output must satisfy, such as "the price comes from the authorised list" or "escalate when the delivery date is missing".
  4. Choose the check. Use an exact comparison where one exists, a rule where it does not, and a written rubric with expert review for judgement calls.
  5. Version it with the code. A case added in March should still run in November.
  6. Rerun on change. New prompt, model, tool, connector or data source: run the set first, read the differences, then decide.

Start small. A modest set that the process owner recognises and trusts is more useful than a large one nobody maintains.

Evals decide how much autonomy each step earns

Autonomy is not a property of "the agent". It is granted per step, and it should follow the evidence. Four levels are enough to reason about most workflows.

The autonomy ladder: four levels from drafts for review to acting within limits, each with the evidence a workflow step needs before it moves up.

Level What the step does Evidence required
1. Drafts for review Prepares a draft that a person finishes Every run leaves a complete trace; reviewers find the drafts usable
2. Proposes, human approves Proposes an action; nothing happens without approval A regression set exists; approvals, edits and rejections are captured
3. Acts, people sample Acts; people review a sample and every exception The regression set keeps passing across changes; edits are rare
4. Acts within limits Acts alone inside defined limits and escalates outside them Limits written down, escalation tested, sampling continues

Two consequences follow. Steps in the same workflow can sit at different levels: reading an incoming order may act within limits while committing to a delivery date still waits for approval. And movement goes both ways. When a regression run fails, or corrections cluster around one type of case, the step returns to approval. If approval steps were designed into the workflow from the start, that is a small change and not a rebuild.

Who owns the regression set?

The company that owns the workflow. A regression set is operational knowledge in executable form: the cases your people handle, the exceptions they know and the answers they accept. It should sit with your custom code, in your environment, under your access rules.

On Nova the split follows the same line as ownership and continuity in our approach. We maintain and license the shared Backbone and reusable modules. Your specific workflow, its cases and its accepted answers are part of your product. That matters on the day you change model, supplier or development team: the set tells the next team what the software must keep doing. It is one of the practical reasons to own your AI software and not only use it.

Where to begin

Define the standard before the build. ScopeRight, an independent assessment and scoping practice, describes how to define evaluation criteria during scoping, and that work carries straight into the design.

Then take three build decisions early: what every step records, where people approve and how their corrections are captured, and where the regression set lives. With those in place, each week of real use makes the next change safer to release.

If you have a workflow where this matters, discuss your build with us.

Key takeaways

  • Evaluation that lasts is a property of the software: every run leaves a trace, and every human decision is kept.
  • Human approval steps are a control and a source of examples at the same time; approvals, edits and rejections state what correct looks like.
  • A regression set built from real corrections is rerun before each change to a prompt, model, tool or data source.
  • Autonomy is granted per workflow step and follows the evidence, in both directions.
  • The regression set is operational knowledge in executable form and should be owned with your custom code.

Frequently asked questions

Does Nova Backbone include an evaluation tool?
No. The Backbone core provides workflow orchestration with approval steps and audit trails, which produce the traces and corrections that evals are built from. The evaluation set-up itself is agreed per project and built as part of your product.
What should a trace contain to be useful for agent evaluation in production?
The input, the context that was retrieved and its source, the tools called, the proposed result, the human decision, and the versions of the prompt, model and configuration. With that, a past run can be replayed against a changed workflow.
Can a workflow step lose autonomy once it has earned it?
Yes. If a regression run fails or corrections cluster around one type of case, the step goes back behind a human approval step until the evidence recovers.
Who owns the regression set when a workflow is built with Nova?
You do. The cases and accepted answers describe your specific business workflow, so they are part of your custom code and stay in your environment. Nova maintains and licenses the shared Backbone and reusable modules.
  • Evaluation
  • Agentic workflows
  • Human oversight
  • Audit trail
  • Regression set
  • Ownership

Have a workflow like this in mind?