Product · Verification

Know the moment your answers change

A golden test set built from your customer's own data, held by us, re-run unattended every time the model underneath you moves - and every regression pinned to the field that caused it. Bought so you are not blindsided.

What you get

Three properties, none of them a feature

Eval tooling can replay a set. What it cannot do is hold the set outside your reach, run it on a schedule you do not control, and refuse a comparison it cannot attribute.

The set is built from your data, and we hold it
Mined from your documents, tickets and past decisions, confirmed once by your expert in under four hours, and versioned from that day on. You can propose changes; changes are logged; a case you failed cannot quietly leave the set.
The runs happen to a fixed schedule, not on request
The full set weekly, a smoke set daily, and a confirm run within the hour of a change we detect - by adapter fingerprint, your CI/CD webhook, provider polling, and behavioural canaries when a provider says nothing at all.
The delta is attributed, or the comparison is refused
Every answer arrives with a fingerprint of what produced it. Exactly one field moved means the delta belongs to that field. Two or more, and the run says confounded rather than blaming the thing it was testing.

The mechanism, in full - what one run does and what crosses the wire

The screen

The morning after a run goes red

A case that passed and then did not, with the one field that moved and the step where the two runs stopped agreeing. Two traces aligned on a byte-identical input is the part that needs somebody holding the same set on both sides of the change.

PresojaProduction Q&A · Insurer A · R-118Windscreen excess, Comprehensive

The model stopped calling check_refund_policy before answering

Both runs retrieved the same two passages and asked for the same three tools on turn 1. On turn 3 the incumbent called check_refund_policy and the candidate did not - it answered from context instead, and dropped the qualifier “once per period of insurance”. This is the shared cause behind 8 of the candidate's 14 newly failing cases.

CaseCV-4471 · v3Tier Aconfirmed by J. Whitfield, claims manager, 12 Febmined from Motor PDS v4.2 §4.2 Windscreen and glassin set v11, sealed 28 Feb, unchanged since

Fingerprint diff

exactly one field moved, so the delta is theirs

1 of 5 moved
generation modelgpt-4-0613gpt-5.6-soldiffers
4 fields identicalsystem prompt 7f3a… · retrieval index v22 · config digest 9c11… · decode params 40de…same

One field. Had two moved, everything below would still render and no verdict would be offered - the panel would read confounded and say which pair to separate.

Execution trace, aligned

one spine, two runs - identical steps collapsed, first divergence expanded

turns 1 and 2 in full - same tools, same order, same two passages retrieved
turn 3 · llmfirst divergence
finish tool_callsasks for check_refund_policy · 1,700 tok in
divergesfinish stopasks for nothing · 1,690 tok in
tool call
check_refund_policy54ms · ok
missingnever calledthis step does not exist in run 847
turn 4 · final answer
784 chars
shorter612 chars −172the missing qualifier is in the difference

Nine steps in one run, eight in the other, and one path on the screen. A trace viewer shows you two lists. This draws the path both runs shared, the step where they stopped sharing it, and asserts that the six collapsed steps agreed.

Both runs are sealed, and this screen is derived from them. Run 812 is entry #41 on Insurer A's chain and run 847 is #48; each manifest carries a hash over the trace above, so a reader can check that what this screen shows is what was recorded on the night.

Illustration with sample data, not a screenshot.

The next step

The numbers are on the page

Two plans, nothing metered, and one question back: whether the price is roughly right.