How Accurate Is RenoCalc's AI, Actually?
Rather than describe RenoCalc's accuracy with adjectives, this page sets out the testing methodology, what's been verified, and — just as importantly — where a builder still needs to check the AI's work before it goes to a client.
Verified The calculation engine has been checked for numerical parity against 27 test fixtures: 132,522 of 132,531 formula cells matched expected output — 99.99% parity. This measures the spreadsheet engine's arithmetic, not the AI's floor-plan reading — see the estimates below for that.
What Gets Tested
RenoCalc's floor-plan AI is evaluated against real, previously-unseen UK residential floor plans covering houses, flats, extensions and loft conversions, spanning architect's CAD drawings, estate agent PDFs, and hand-drawn sketches. For each test plan, the AI's output is compared against a human-measured ground truth across:
- Room identification — is each room detected and correctly classified?
- Room area and perimeter variance against measured ground truth
- Door and window opening detection
- Downstream quantity and material-cost variance once pricing is applied
- Proportion of plans requiring manual calibration or room correction
- End-to-end processing time from upload to reviewed quote
Results
| FactSample size | 27 test cases (5 real-world, 22 synthetic); expanded validation in progress. Counted, not estimated. |
| EstimateRoom identification | ~90% on clean, labelled architectural plans; 75–85% on hand-drawn or unlabelled sketches. Room type classification is weaker than room detection. |
| EstimateArea / perimeter variance | ±10% overall (perimeter ±3–5%, area ±6–10%). Scale is anchored to a user-supplied known dimension, which removes systematic scale error; area compounds two linear readings so its error runs roughly double the linear figure. |
| Estimate — weakest metricOpening (door/window) detection | ~85% for count detection (is a door/window there at all). Dimension extraction is materially weaker — e.g. a door correctly counted can still return a width/height of 0. Treat this as "openings found," not "openings measured," until dimension extraction is validated separately. |
| EstimateOverall estimate variance vs. actual cost | ±20%. Dominated by material/labour rate and scope assumptions, not measurement error — broadly comparable to a RICS order-of-cost estimate at pre-site stage. |
| Fact + EstimateManual calibration/correction | Scale calibration is mandatory and enforced in code — 100% of quotes require it, by design. Beyond calibration, an estimated ~40% of cases need at least one further room or dimension correction before the quote is finalised. |
| EstimateAverage processing time | ~90 seconds typical machine time (≈40–90s AI vision call + ≈15s spreadsheet generation). "Under 3 minutes" elsewhere on the site is a safe ceiling / client timeout figure, not the typical case. |
Read together: the calculation engine's arithmetic is measured and verified. The AI's floor-plan reading is estimated from a small internal test set and is honest about where it's weakest — opening dimensions in particular are not yet reliable, which is exactly why every detected room and dimension is editable before a quote is generated. These estimates will be replaced with measured figures as the validation dataset grows past its current 27 fixtures.
What Affects Accuracy
Accuracy depends on factors including drawing quality, drawing scale, legibility, completeness of the drawing, hidden construction conditions, and the accuracy of user-supplied project information. Where a drawing is unscaled or its scale cannot be reliably inferred, the user calibrates it against a known dimension before quantities are generated.
Every measurement, rate, quantity and assumption in the output spreadsheet can be edited before a quotation is finalised — this is the mechanism by which a builder corrects for anything the AI got wrong before a client sees the number.
Limitations
RenoCalc is an estimating assistance system, not a replacement for a site survey, structural engineering, building control approval, or professional judgement on complex or non-standard projects. It does not detect conditions hidden behind walls, floors or ceilings — damp, asbestos, structural defects, or buried services. Full detail: /official-information/.
Update History
| 14 September 2026 | Page published. Methodology and testing scope documented. |
| 14 September 2026 | Added verified calculation-engine parity result (99.99%, 27 fixtures) and engineering estimates for AI vision metrics (room identification, area/perimeter variance, opening detection, cost variance, calibration rate, processing time), each labelled Fact or Estimate. Opening-dimension detection flagged as the weakest current metric. |
This page is reviewed monthly. Next scheduled review: October 2026. Estimates are updated to measured figures as the validation dataset grows beyond its current 27 fixtures.