Skip to content
AI

AI Agent Evaluation: Grade the Database, Not the Reply

AI Agent Evaluation: Grade the Database, Not the Reply

Short version: if your agent writes records, your test should read records. Whatever the agent says in its final message is a claim. The database is the evidence.

That sentence is easy to nod along to and easy to ignore, so here’s the example that made me stop ignoring it. Microsoft and Hugging Face just published a post about a benchmark called ThinkingBox, and it opens with a customer-support agent that does nine careful tool calls on a delayed $745 appliance. It reads the order, checks tracking, searches the refund policy twice, opens a ticket, documents the timeline. Then it closes the ticket as solved and asks the customer if there’s anything else it can help with. The carrier exception was still open, so the ticket should have been on hold. A grader reading the transcript would have seen nine well-formed tool calls and a polite sign-off. One database field disagreed.

I’ve been building small agents that touch real records for clients, mostly through n8n and a few custom services, and that story is uncomfortably familiar. This post is about how I’d test for it, using the ThinkingBox write-up as the evidence and a boring Postgres setup as the tooling. I haven’t run the benchmark itself. Its setup involves a Typesense instance, several MCP servers and an OpenEnv server, and I’m not going to pretend I did that over a weekend.

What the benchmark actually measured

ThinkingBox runs agents against 507 stateful business workflows across five domains (retail, auto insurance, travel, neobank and consulting). Each task runs 20 times from an identical clean backend. Grading looks at the terminal state of the backend and the side effects, not at the agent’s words.

The numbers that stuck with me come from the authors’ ablation across 121,680 valid trials on 12 models. Of those, 79,853 attempts failed the executable checks. And 67.24% of the failures still ended cleanly, called a state-changing tool and reported no final tool error. Within those quiet failures, the checks found wrong field values in 77.61%, unintended extra effects in 43.30% and missing required effects in 25.36%. Those categories overlap, so don’t add them up.

Read that first figure again. Two out of three failures looked like successes from the outside. If your monitoring is “did the agent finish without an exception”, you are blind to most of the failures this benchmark found.

I’ll add the skepticism the post doesn’t supply itself. This is a vendor-authored benchmark on synthetic workflows, and the cost figures are list-price estimates the authors computed, not invoices. The absolute pass rates won’t transfer to your support bot. The shape of the finding is what transfers: agents fail quietly, and the evidence is in the records.

The consistency problem is a separate problem

The second half of the post is about repetition, and it’s the part I’d tape to a monitor. A model that gets a refund right once and wrong four times out of five is not a refund agent. So the authors report pass@1, pass@20 and a literal count of tasks that passed all 20 attempts.

The model comparison is where it gets uncomfortable. Claude Opus 5.5 scores 67.16% pass@1 against 66.50% for Claude Opus 5, and solves more tasks at least once. But both pass exactly 241 tasks on every one of 20 attempts. Half a point of headline accuracy bought zero extra dependability. On the other side, Kimi-K3 solves 476 of 507 tasks at least once, the broadest coverage in the field, yet only 68 of them every time.

Then the money. The cheapest model per successful attempt, GPT-5.6 Sol at $0.127, comes out at $9.76 per dependable task, because only 82 of its tasks pass all 20 runs. GPT-6 Astra sits at $7.45 per dependable task and Opus 5.5 at $7.80. The cheapest way to get one right answer isn’t the cheapest way to get a reliable one.

I wrote about the same shape of problem in my earlier post on agent testing and the repeat-run gap. The ThinkingBox data is a much bigger sample making the same point, and I’m happy to be outgunned there.

A state-check harness you can build this week

You don’t need 507 tasks. You need one workflow, a seeded database and a loop. Postgres makes the reset step nearly free, because CREATE DATABASE ... TEMPLATE copies a whole database. The docs on template databases explain the mechanics, with one restriction worth remembering: nobody else can be connected to the template while you copy it.

The plan has four parts. Seed a template database with fixture rows. For each run, clone it into a scratch database. Point your agent’s tools at the scratch database and give it the task. Then assert on the rows with plain SQL.

import psycopg

def fresh_db(admin, run_id):
    name = f"agent_run_{run_id}"
    admin.execute(f"CREATE DATABASE {name} TEMPLATE agent_fixture")
    return name

def check_ticket_on_hold(conn, ticket_id):
    row = conn.execute(
        "SELECT status FROM tickets WHERE id = %s", (ticket_id,)
    ).fetchone()
    assert row[0] == "hold", f"status was {row[0]!r}, expected 'hold'"

That one assertion is the ThinkingBox failure from the intro, written in six lines. The agent’s transcript never enters the picture.

Wrong values are only one of three failure kinds

The benchmark’s breakdown tells you what else to assert. Wrong field values were the most common finding, but extra effects showed up in 43.30% of those quiet failures and missing effects in 25.36%. A test that only checks the field you expected to change catches the first kind and misses the second.

For extra effects, snapshot everything the agent shouldn’t touch before and after, then diff. A cheap version hashes each table you care about:

def table_fingerprints(conn, tables):
    out = {}
    for t in tables:
        out[t] = conn.execute(
            f"SELECT md5(string_agg(t::text, ',' ORDER BY t::text)) FROM {t} t"
        ).fetchone()[0]
    return out

Run it before the agent starts and after it finishes. Any table outside the allowed set whose fingerprint changed is an unintended effect, and the test should fail with the table name. For missing effects, assert the existence of the rows the workflow must create, such as the audit note or the refund record, not only the status flip.

Then wrap the whole thing in a loop of 20. I’d start with 10 if the model calls are expensive. Count how many runs pass every assertion. That single integer, passes out of N, is the number I’d put on a dashboard, not an average score.

Where the failures landed

The authors assign each failed trace one diagnostic signature. About 80% of failures, in their unweighted average, were tool usage. Wrong state updates were about 10%, incomplete resolutions about 7%, and failing to take any state-changing action about 3%. The post’s own summary is that agents usually get far enough to try the workflow, then fail to recover from tool errors, failed preconditions or empty lookups.

That matches what I see in n8n flows, where the bug is rarely the model’s reasoning and almost always what happens when a lookup returns nothing. If you want a concrete case, my post on the n8n bug that cost a client eleven days is about exactly that kind of silent empty result. So put empty-lookup fixtures in your seed data on purpose: an order that doesn’t exist, a customer with two records, a ticket that is already closed. Those are the cases where the polite sign-off hides the damage.

What I’d do differently from the benchmark

Two caveats about my own approach. A clone-per-run harness tests your agent on your fixtures, and fixtures are written by the person who wrote the agent, which means they flatter it. Have someone else add the awkward rows. And assertions on final state say nothing about whether the customer got a useful answer. The ThinkingBox example had that problem too, since the customer never got a real answer to her question. A database check can’t see it. I’d still add a human spot check on a sample of transcripts, because the two methods fail on different things.

I build this kind of agent plumbing for clients, and you can see what I work on here.

This week, pick the one workflow in your system where an agent writes to a record. Write the single SQL assertion for its correct end state, run the agent against a cloned database ten times, and write down how many runs pass. If the answer is ten, add an awkward fixture and run it again.