Accuracy · the validation record

How wrong are we? Here is the record.

Our insurance figures are estimates based on public data — never quotes. Since mid-July we have been gathering real quotes for fixed driver profiles and scoring the estimates against them, sitting by sitting. This page publishes what came back, including the two worst numbers we have produced.

How the testing works

Every insurance figure on Vouched prices one deliberately narrow profile: a new driver with no no-claims history on a comprehensive black-box policy, anchored to London and scaled from there. To test it we gather real quotes for that same profile on a comparison site, one car and age at a time, and score two separate things — the level (how far the pounds are out) and the ranking (whether the cars come out in the right order as a shortlist).

Those are different failures with different fixes, which is most of what the record below is about. A level miss is one number applied to everything. A ranking miss means a car is in the wrong place on your shortlist, and no single correction can fix it.

What the estimates say today

We price 164 cars. 48 have been quoted directly, or sit in a family that has; 87 never have. Each figure carries that distinction as a confidence flag, and the range we display around it widens with the flag — the less we have checked a car against real quotes, the more room we leave.

ConfidenceRange shown around our figureCarsWhat that means
High35% below to 25% above48Quoted directly at up to four ages, or in a family that was — and shown to someone whose postcode we know.
Medium35% below to 35% above29Priced from published data for its family — or a quote-validated car shown without a postcode, which is a London figure applied blind.
Low35% below to 50% above87Never quote-validated. The wide band is deliberate: the misses on 27 July were on exactly this kind of car.

Those ranges are read straight off the code that draws them, so this table cannot drift from what the site actually shows you — including the extra width the market notice below adds while it is live.

Heads-up: Market update (27 July 2026): two major insurer groups — including the cheapest black-box provider in most of our data — left the comparison panel we calibrate against, and cheapest quotes for young drivers rose roughly 50%. Our estimates now track this new, higher market. If those insurers return, real quotes could come in below the range shown — get a real quote before you budget.

The record, sitting by sitting

Each entry is a real-quote sitting: what we gathered, what it showed, and what changed in the model because of it. Figures are frozen as measured — they are history, not claims about today — and each one names the section of our internal validation record it is copied from.

14 July 2026 · Market before 27 July 2026

The first real-quote sample

32 real quotes — 4 cars at 4 ages (17, 19, 21, 24), each with the cheapest black-box and the cheapest non-black-box policy, from one London postcode on one comparison journey.

Median absolute error, model as shipped
17.9% §4
Cells inside ±25%, as shipped
10 of 12 §4
Median absolute error after re-anchoring
10.8% §4
Ranking agreement with the real quotes
0.85 mean Spearman §4
What turning down a black box cost a 17-year-old
×1.76 §5

What it showed. Eleven of the twelve core cells came in under the cheapest real quote. A one-sided miss like that is a level problem, not noise — and the age curve was flatter than the market past 21.

What changed. The level was re-anchored ×1.168 and the age-24 point on the curve re-pinned from 0.578 to 0.656.

Read this with the caveat: Every winning quote in this sample came from one insurer group, in one postcode, on one day — so the level it produced carried single-panel risk from the day it shipped. That risk is exactly what came due on 27 July.

24 July 2026 · Market before 27 July 2026

The age-curve sitting

122 real quotes across 61 car-age cells and 25 cars, insurance groups 2 to 30, one fixed driver profile on one comparison site.

Median absolute error before the re-pin
19.8% §7
Cells inside ±25% before
36 of 61 §7
Median absolute error after the re-pin
15.2% §7
Cells inside ±25% after
47 of 61 §7
Ranking agreement — a fail against our 0.80 bar
0.75 mean Spearman §7

What it showed. The overall level held; the shape was wrong. Cheap cars were under-estimated and expensive ones over-estimated, and measured against the quotes themselves the insurance-group slope was 0.0208–0.0241 per group — about half the 0.0433 we had taken from a published table. Published group contrasts overstate what insurers actually quote.

What changed. Group slope re-pinned 0.0433 → 0.022, and the level moved ×1.168 → ×1.406.

Read this with the caveat: The ranking failures were re-orderings — hybrids quoted above their group, EVs and executive cars below — and no choice of slope can produce those. Fixing them needed per-car quote evidence, not another level correction.

25 July 2026 · Market before 27 July 2026

Per-car quote blending

No new quotes. Each of the 25 cars already quoted took a correction factor from its own observed cells, blended half-and-half with the published-data model.

Median absolute error on the 61-cell sample
15.2% → 11.1% §8
Cells inside ±25%
47 → 57 of 61 §8
Ranking agreement
0.75 → 0.93 §8
Ranking agreement at age 19
0.42 → 0.89 §8

What it showed. Correcting per car, rather than per insurance group, recovered the ordering the slope work could not. Age 19 — the widest cross-section and the worst-ranking age — moved furthest.

What changed. Those 25 cars now ship a quote-blended premium and the high confidence flag; every other car keeps the published-data model.

Read this with the caveat: Read these as in-sample, because they are: the sample that supplied the corrections is the sample they are scored against, so a low error there is partly definitional rather than predictive. The honest test is the next independent sitting — never read this as current accuracy.

27 July 2026 · Market before 27 July 2026

The first coverage sweep — and the panel change

9 small SUVs and crossovers at 19, none of them quote-blended: a genuinely out-of-sample test of the method on cars it had never seen.

Median absolute error, out of sample
8.3% §9
Cells inside ±25%
6 of 9 §9
Ranking agreement — a real, structural fail
0.75 mean Spearman §9
Worst miss — Toyota C-HR, a hybrid
−27% §9
Worst miss — MG ZS EV
−28.9% §9
Control car at 19, before and after the panel change
£1,269.87 → £1,898.67 (×1.50) §9
The same control car at 21
£1,083 → £1,734.76 (×1.60) §9

What it showed. On level it held out of sample. On ordering it failed structurally, and both worst misses were powertrain cases — the model under-prices hybrids, and cheap older EVs run the opposite way to the EVs the blending was trained on. Then, mid-sitting, two insurer groups stopped appearing on the comparison panel we calibrate against and did not come back. One of them was the cheapest black-box provider in 56 of the 61 cells the level had been fitted to.

What changed. Further sittings were frozen, a dated market notice went onto every figure we show, and the level re-pin started the same day.

Read this with the caveat: That out-of-sample result scored a comparison panel that no longer exists. It is evidence the method generalises to cars we had never quoted — it is not a claim about today's accuracy, and we do not present it as one.

28 July 2026 · Market from 27 July 2026

Re-pinned to the market that exists now

Six control quotes gathered after the panel change, four of them for cars that had never been quote-blended — so the new level was checked against cars it was not fitted from.

Residual spread across the six control cells
−7.1% to +9.0% §9.1
Control cells never used to fit the model
4 of 6 §9.1

What it showed. One uniform step, derived at build time from those six control cells and applied to every car in the set, brought the estimates back onto the market as it now trades.

What changed. Estimates track the panel as it is today, and the low end of every range widened: if those insurers return, a real quote could land below the range we show.

Read this with the caveat: The 1.0% median error across those six cells is 1.000 by construction — the step is fitted to make that median exactly 1, so it is not an independent pass mark the way the sittings above are. The spread is the honest read.

Section marks point at our validation record — the working document these figures are copied from, kept in the repository that builds this site, alongside the raw quote data and the scoring script.

What we do not claim

  • None of it is a promise about your quote. These are scores against fixed test profiles on particular days. Your age, address, car and history move the number, which is why every estimate on the site points you at a real quote before you budget.
  • We never add the two markets together. Two insurer groups left the comparison panel we calibrate against on 27 July. Everything measured before that date scored a market that no longer exists, so those figures stay stamped with their era and are never averaged into a single headline. We now watch that market directly — panel watch publishes each sitting's cheapest quote and which insurers we saw, so the next change is visible as it happens.
  • Scores measured on the sample that trained a correction are labelled as such. A model checked against the data it was fitted to will always flatter itself. Where that is true above, the caveat is printed with the figure.
  • We publish the misses by name. The two worst — a hybrid and a cheap electric car, both under-priced by more than a quarter — are in the record above with the cars named. They are the reason the low-confidence range is as wide as it is.

Whether that panel change is permanent is still unresolved — two days of control quotes cannot tell us. The next test is a persistence check at age 24 in early August. If the insurers who left come back, we re-pin again, from a new frozen sample, and it gets published here.

The part you can change

The next honest number on this page comes from quotes we did not gather ourselves. When enough have arrived to compare like with like — same market, same shape of profile — the gap between our estimates and what people are really quoted gets published here, misses included, the same way everything above is.

If you have been quoted recently, that takes half a minute. Nothing identifying is stored, and no email is needed.

Got a real quote? Tell us what it was.

Thirty seconds, no email needed — every real quote sharpens the estimates for the next person.

Stored anonymously — quote amount, age and area only. Used solely to improve the estimates.

Vouched does not give insurance advice; where insurance is involved we introduce you to authorised providers. Our figures are indicative estimates based on cited public data, never quotes — see how it works for what goes into them.