PROOF

We publish how we are wrong.

We are pre-reference-customer, so there are no case studies on this page yet. What we can show you is the machinery that will produce them — and that already runs against every prediction the platform makes.

MEASURED, NOT CLAIMED

Every prediction is scored against what actually happened.

A prediction is written down with a due date. When that date arrives, the platform compares it to the measured outcome and records the absolute error and whether it was a hit. The tolerances are fixed in code and unit-tested — we cannot quietly widen them to flatter a result.

Every prediction is scored against what actually happened.
PredictionCounts as a hit withinMeasured against
Schedule delay±7 daysForecast delay at the horizon date
Cost forecast±10%Forecast amount at the horizon date
Risk level±20 pointsProject risk score at the horizon date

Hit rate and mean absolute error are exposed to every customer through the platform’s own analytics, per organization. When we have enough scored predictions across live projects to be statistically meaningful, the aggregate lands on this page — good or bad.

ENFORCED IN CODE

The evidence policy is a constraint, not a slogan.

These four behaviours are implemented in the reasoning service and covered by the release test suite. They are not prompt instructions, which a model is free to ignore.

01

No record, no confident answer

When nothing in the project matches the question, the response is flagged provisional, confidence is capped at 0.4, and the answer text states it must not be used as the basis for a contractual decision.

02

The AI’s own citations are verified

Every reference in an answer is checked against the records actually retrieved. If the model cites something that was not, confidence is capped at 0.45 and the answer names the unverified reference.

03

Provenance on every response

Each answer reports whether it was model-backed, which provider and model produced it, the retrieval method, and the schedule sample size it reasoned over.

04

Thin samples refuse to look confident

With fewer than three measured activities the forecast returns the recorded baseline delay and a warning, instead of a distribution that would look authoritative and be meaningless.

RELEASE VALIDATION

What was actually tested in the last release.

These are the numbers from the product’s own validation report, not a marketing summary of it.

31Automated tests passing against an isolated database per run
29Live end-to-end checks against a running API and asset worker
100Pilot readiness score from the end-to-end validation chain
143API endpoints across the Construction OS service surface
BOUNDARIES

We publish our limitations.

The product ships a numbered list of explicit boundaries — uncalibrated risk heuristics, lexical rather than semantic retrieval, forecast that does not traverse the dependency network. Ask for it during evaluation and we will hand it over before you ask twice.

YOUR PROJECT

Be the first reference.

Pick one project and one decision that is currently hard to make. We agree the success test in writing before we start, and we report against it honestly — including when it fails.

Start an Enterprise PilotBook a Demo
Proof & Accuracy | OneAI Construction