Field note

What flat scores hide.

Why flat aggregate scores can hide changed AI agent behavior, affected workflows, and evidence gaps.

For AI platform and evaluation teamsPublished August 7, 2026

The dangerous release is not always the one that lowers the headline score. Sometimes the score stays flat because an improvement in the common path cancels out a regression in a smaller, higher-consequence path.

Aggregate metrics can be reported, but they are not evidence of behavior. The problem starts when a summary is treated as an explanation of every path the agent can take.

The pattern

Aggregate metrics compress many decisions into one number. When behavior migrates between workflows, tools, cohorts, or failure modes, the number can remain stable while the agent’s practical decision rule changes.

What to inspect instead

Look for changed decisions, changed influence, affected workflows, and evidence coverage. The strongest review leads with behavior diffs and representative traces, then records what the evaluation did not exercise.

Start with the largest behavior changes, not only the largest numerical deltas. A small shift in a high-consequence workflow can deserve more attention than a broad improvement in a low-risk path.

Why the refund example matters

In the refund-agent example, pressure cues become more influential while verified evidence becomes less influential. The story is visible in the behavior diff even though the overall score barely moves.

Read next

Use AI agent evaluation, then read the guide to AI agent regression testing or inspect the refund-agent behavior diff.