No record, no confident answer
When nothing in the project matches the question, the response is flagged provisional, confidence is capped at 0.4, and the answer text states it must not be used as the basis for a contractual decision.
We are pre-reference-customer, so there are no case studies on this page yet. What we can show you is the machinery that will produce them — and that already runs against every prediction the platform makes.
A prediction is written down with a due date. When that date arrives, the platform compares it to the measured outcome and records the absolute error and whether it was a hit. The tolerances are fixed in code and unit-tested — we cannot quietly widen them to flatter a result.
| Prediction | Counts as a hit within | Measured against |
|---|---|---|
| Schedule delay | ±7 days | Forecast delay at the horizon date |
| Cost forecast | ±10% | Forecast amount at the horizon date |
| Risk level | ±20 points | Project risk score at the horizon date |
Hit rate and mean absolute error are exposed to every customer through the platform’s own analytics, per organization. When we have enough scored predictions across live projects to be statistically meaningful, the aggregate lands on this page — good or bad.
These four behaviours are implemented in the reasoning service and covered by the release test suite. They are not prompt instructions, which a model is free to ignore.
When nothing in the project matches the question, the response is flagged provisional, confidence is capped at 0.4, and the answer text states it must not be used as the basis for a contractual decision.
Every reference in an answer is checked against the records actually retrieved. If the model cites something that was not, confidence is capped at 0.45 and the answer names the unverified reference.
Each answer reports whether it was model-backed, which provider and model produced it, the retrieval method, and the schedule sample size it reasoned over.
With fewer than three measured activities the forecast returns the recorded baseline delay and a warning, instead of a distribution that would look authoritative and be meaningless.
These are the numbers from the product’s own validation report, not a marketing summary of it.
The product ships a numbered list of explicit boundaries — uncalibrated risk heuristics, lexical rather than semantic retrieval, forecast that does not traverse the dependency network. Ask for it during evaluation and we will hand it over before you ask twice.
Pick one project and one decision that is currently hard to make. We agree the success test in writing before we start, and we report against it honestly — including when it fails.