How the proof was measured

Numbers your client can question—and your agency can defend.

Every claim traces to a fixed comparison, live repair, or dated operating record. Inspect the conditions behind each result, then use only the claim the opportunity supports.
Second dated lift
+2.85
First dated lift
+0.75
Prompts improved in best run
10/10
Real-browser repair checks
20/20

Same-contract quality observations

Two dated results under one unchanged contract.

Select either run to inspect its result. The observations are deliberately unconnected: two points in a short window do not establish a trend or cause.
+0.75first dated lift
+2.85second dated lift
103 minbetween run starts
+3.00

Provider, model, ten-item test set, scoring method, and run settings stayed fixed

Run 2 · 10:36:24 PM UTC

+2.85 mean lift

Prompts with higher scores
10 / 10
Aria mean
12.45 / 14
Baseline mean
9.6 / 14
Recorded losses
0
Dated source run · 2026-05-18T22-36-24-995Z
Inspect the recovery observation

One exact test did not satisfy the fixed comparison contract and was excluded. The next verified observation returned positive in 3m 46s; the strongest reached +3 in 6m 31s.

Selected complete observation+1 lift

226 seconds after the excluded observation · same fixed test · verified result

01 · Same-contract observations

Two dated results. One unchanged measurement contract.

Two complete runs used the same provider, model, ten-item test set, scoring method, and run settings. They are shown as separate observations so a short window cannot masquerade as a trend.

The first run observed +0.75 mean lift. The second observed +2.85, 103 minutes later.Your agency can defend a quality claim without asking the buyer to trust a changed test.
  • Prompt wins moved from 6 of 10 to 10 of 10
  • Mean lift measured +0.75 in the first run and +2.85 in the second
  • Two dated observations only; no causal, trend, or forecast claim
02 · Pre-execution challenge

Every counted challenge names what would prove it wrong.

The challenge record counts checks, not vague claims that work was reviewed. Each counted check is attached to an action and names a concrete condition that would prove the move unsafe or wrong.

123,777 structured checks covered 8,275 unique actions.Plausible wrong moves can be challenged before they become client-facing rescue work.
  • Every counted check names what would prove the move unsafe or wrong
  • Observed from June 5, 2026 through July 13, 2026
  • Multiple checks can examine one action from different failure angles
03 · Expanded comparison

The advantage held when the prompt set doubled.

A later 20-prompt run kept the same-model comparison visible as a separate contract. It is not appended to the ten-prompt line, so expansion cannot masquerade as a trend.

Aria produced stronger operating judgment on 16 of 20 prompts, with a mean lift of +1.675 points out of 14.The same provider can deliver stronger operating judgment without a model migration.
  • Baseline mean · 9.2 / 14
  • Aria mean · 10.875 / 14
  • Warnings · 14 baseline vs 6 with Aria
04 · Repair lineage

A visible failure became protection the next client can inherit.

The incident proof begins with one observed interface regression, follows the repair into source, then exercises the corrected behavior in the browser. Reuse begins only after the real interaction passes.

One incident became 5 regression protections; 20 of 20 browser checks passed with zero browser or console errors.A repaired failure becomes protection the next client does not have to fund again.
  • Repair lineage · cdd7a72d1
  • Pointer, scrolling, overlay, and continuation behavior were exercised
  • The public record omits private runtime and session identifiers
05 · Learning promotion

Memory is cheap. Earned reuse is the advantage.

Thousands of observations can enter the learning system without becoming policy. The record separates learning rows, candidates, canary results, and accepted promotions so only supported improvements shape future work.

4,855 learning records were examined; 11 improvements earned durable reuse.Learning becomes reusable delivery capacity only after it earns promotion.
  • 4,855 unique correction records
  • 4,855 complete correction contracts
  • 4,845 evidence references
  • Promotions reached tests, skills, owner guidance, routing, and a control boundary

Claim discipline

Powerful enough to inspect. Precise enough to trust.

  • The quality record contains two complete, comparable observations—not a trend, causal claim, or statistical forecast.
  • A run that did not satisfy the fixed comparison contract is excluded rather than plotted as a quality result.
  • The later 20-prompt comparison is a separate expanded contract and is not connected to the ten-prompt series.
  • The broader three-task run had two fetch failures; only the completed TypeScript workload appears in the commercial comparison.
  • Live provider output is stochastic; these dated observations do not guarantee the same score on future prompts.

Apply this advantage to one live opportunity

Bring the deal. Leave with the result your client has to feel.

Request the private fit decision
Machine-readable appendix

The buyer-facing explanation above is bound to sanitized, dated records. Analysts can inspect the same-contract observations and the broader public operating record directly.

Open dated observations Open public record