Illustrative score benchmarks
Early, illustrative checks of the current scoring model (v2.3) against a mixed set of benchmark archetypes and a small number of measured/monitored homes. This page shows where assumptions still dominate — it is not a measured MAE/ranking validation report, not ScoreLab promote evidence, and not a compliance certificate.
Known limitations
- Measured sample size is still small — most rows are benchmark archetypes.
- Confidence is estimated from input completeness, not calibrated prediction error.
- Architect design-brief quantities are separate indicative bands — not SAP/PHPP outputs.
| Archetype | EPC band | Indicative score band | Source |
|---|---|---|---|
| Victorian mid-terrace (uninsulated) | D–E | 38–48 | benchmark |
| 1980s cavity semi (partial loft fill) | D | 45–55 | benchmark |
| Modern detached (Building Regs 2013+) | B–C | 58–68 | benchmark |
| Deep retrofit (fabric + ASHP) | B | 72–82 | benchmark |
| Monitored pilot home (anonymised) | C | 52–58 | measured |
We do not publish MAE/RMSE or ranking metrics until a labelled, non-overlapping measured sample and real eval path exist. Until then, use this illustrative table and the Passport validation lab for audience-specific projection tests — not as proof of model accuracy.