Run 2 · 10:36:24 PM UTC
+2.85 mean lift
- Prompts with higher scores
- 10 / 10
- Aria mean
- 12.45 / 14
- Baseline mean
- 9.6 / 14
- Recorded losses
- 0
How the proof was measured
Same-contract quality observations
Provider, model, ten-item test set, scoring method, and run settings stayed fixed
Run 2 · 10:36:24 PM UTC
One exact test did not satisfy the fixed comparison contract and was excluded. The next verified observation returned positive in 3m 46s; the strongest reached +3 in 6m 31s.
226 seconds after the excluded observation · same fixed test · verified result
Two complete runs used the same provider, model, ten-item test set, scoring method, and run settings. They are shown as separate observations so a short window cannot masquerade as a trend.
The first run observed +0.75 mean lift. The second observed +2.85, 103 minutes later.Your agency can defend a quality claim without asking the buyer to trust a changed test.The challenge record counts checks, not vague claims that work was reviewed. Each counted check is attached to an action and names a concrete condition that would prove the move unsafe or wrong.
123,777 structured checks covered 8,275 unique actions.Plausible wrong moves can be challenged before they become client-facing rescue work.A later 20-prompt run kept the same-model comparison visible as a separate contract. It is not appended to the ten-prompt line, so expansion cannot masquerade as a trend.
Aria produced stronger operating judgment on 16 of 20 prompts, with a mean lift of +1.675 points out of 14.The same provider can deliver stronger operating judgment without a model migration.The incident proof begins with one observed interface regression, follows the repair into source, then exercises the corrected behavior in the browser. Reuse begins only after the real interaction passes.
One incident became 5 regression protections; 20 of 20 browser checks passed with zero browser or console errors.A repaired failure becomes protection the next client does not have to fund again.Thousands of observations can enter the learning system without becoming policy. The record separates learning rows, candidates, canary results, and accepted promotions so only supported improvements shape future work.
4,855 learning records were examined; 11 improvements earned durable reuse.Learning becomes reusable delivery capacity only after it earns promotion.Claim discipline
Apply this advantage to one live opportunity
The buyer-facing explanation above is bound to sanitized, dated records. Analysts can inspect the same-contract observations and the broader public operating record directly.
Open dated observations Open public record