Accuracy & Testing Methodology — RenoCalc
Try Free Now
Accuracy & Testing

How Accurate Is RenoCalc's AI, Actually?

Rather than describe RenoCalc's accuracy with adjectives, this page sets out the testing methodology, what's been verified, and — just as importantly — where a builder still needs to check the AI's work before it goes to a client.

Last updated: 14 September 2026  ·  See also: /official-information/

Status: one verified engine result, plus engineering estimates pending full field validation.

The calculation engine's numerical accuracy has been directly measured and is reported below as a verified result. The AI vision metrics (room identification, dimension variance, opening detection) are currently engineering estimates — informed by the model's behaviour on a small internal test set (27 fixtures: 5 real-world, 22 synthetic), not yet a full measured field study. They are labelled as estimates throughout this page and will be replaced with measured figures as the validation dataset grows. Reviewed monthly — see Update History below.

Verified The calculation engine has been checked for numerical parity against 27 test fixtures: 132,522 of 132,531 formula cells matched expected output — 99.99% parity. This measures the spreadsheet engine's arithmetic, not the AI's floor-plan reading — see the estimates below for that.

What Gets Tested

RenoCalc's floor-plan AI is evaluated against real, previously-unseen UK residential floor plans covering houses, flats, extensions and loft conversions, spanning architect's CAD drawings, estate agent PDFs, and hand-drawn sketches. For each test plan, the AI's output is compared against a human-measured ground truth across:

  • Room identification — is each room detected and correctly classified?
  • Room area and perimeter variance against measured ground truth
  • Door and window opening detection
  • Downstream quantity and material-cost variance once pricing is applied
  • Proportion of plans requiring manual calibration or room correction
  • End-to-end processing time from upload to reviewed quote

Results

FactSample size27 test cases (5 real-world, 22 synthetic); expanded validation in progress. Counted, not estimated.
EstimateRoom identification~90% on clean, labelled architectural plans; 75–85% on hand-drawn or unlabelled sketches. Room type classification is weaker than room detection.
EstimateArea / perimeter variance±10% overall (perimeter ±3–5%, area ±6–10%). Scale is anchored to a user-supplied known dimension, which removes systematic scale error; area compounds two linear readings so its error runs roughly double the linear figure.
Estimate — weakest metricOpening (door/window) detection~85% for count detection (is a door/window there at all). Dimension extraction is materially weaker — e.g. a door correctly counted can still return a width/height of 0. Treat this as "openings found," not "openings measured," until dimension extraction is validated separately.
EstimateOverall estimate variance vs. actual cost±20%. Dominated by material/labour rate and scope assumptions, not measurement error — broadly comparable to a RICS order-of-cost estimate at pre-site stage.
Fact + EstimateManual calibration/correctionScale calibration is mandatory and enforced in code — 100% of quotes require it, by design. Beyond calibration, an estimated ~40% of cases need at least one further room or dimension correction before the quote is finalised.
EstimateAverage processing time~90 seconds typical machine time (≈40–90s AI vision call + ≈15s spreadsheet generation). "Under 3 minutes" elsewhere on the site is a safe ceiling / client timeout figure, not the typical case.

Read together: the calculation engine's arithmetic is measured and verified. The AI's floor-plan reading is estimated from a small internal test set and is honest about where it's weakest — opening dimensions in particular are not yet reliable, which is exactly why every detected room and dimension is editable before a quote is generated. These estimates will be replaced with measured figures as the validation dataset grows past its current 27 fixtures.

What Affects Accuracy

Accuracy depends on factors including drawing quality, drawing scale, legibility, completeness of the drawing, hidden construction conditions, and the accuracy of user-supplied project information. Where a drawing is unscaled or its scale cannot be reliably inferred, the user calibrates it against a known dimension before quantities are generated.

Every measurement, rate, quantity and assumption in the output spreadsheet can be edited before a quotation is finalised — this is the mechanism by which a builder corrects for anything the AI got wrong before a client sees the number.

Limitations

RenoCalc is an estimating assistance system, not a replacement for a site survey, structural engineering, building control approval, or professional judgement on complex or non-standard projects. It does not detect conditions hidden behind walls, floors or ceilings — damp, asbestos, structural defects, or buried services. Full detail: /official-information/.

Update History

14 September 2026Page published. Methodology and testing scope documented.
14 September 2026Added verified calculation-engine parity result (99.99%, 27 fixtures) and engineering estimates for AI vision metrics (room identification, area/perimeter variance, opening detection, cost variance, calibration rate, processing time), each labelled Fact or Estimate. Opening-dimension detection flagged as the weakest current metric.

This page is reviewed monthly. Next scheduled review: October 2026. Estimates are updated to measured figures as the validation dataset grows beyond its current 27 fixtures.

Try It on Your Own Floor Plan

The best way to judge accuracy is against a plan you already know the real measurements for. First quote is free.

Try RenoCalc Free View Official Information