The five-fact test, applied to a real retrieval system with real users, and the measurements it produced. Including the question it still gets wrong.
The recorded briefing hands you a standard for judging any AI system. A standard nobody has applied to their own work is a slogan, so this is that standard turned around and pointed at a system we built.
The subject is a search and retrieval tool over one authoritative text, the DSM-5-TR. It is not a product and nothing here is for sale. The shape of the problem is the one an associate has with a practice guide or a code: a bounded authoritative text, sections and page numbers, and a searcher who knows what they need but not what it is called or where it lives.
Retrieval is the first objective, not a smaller version of the goal. Everything a language model could add on top inherits whatever this layer hands it. Most of the work is calibration, none of it demos well, and it is where the trust actually comes from.
[DSM-5-TR | p. 157 | Section II]. The passage is the answer, not
a footnote attached to one.The minimum support score was set to -6. That number was calibrated
against a previous scoring model whose output ran from about -4.6 to +8.9. The
replacement model scores from 0 to 1.
So the threshold was not badly tuned. It was below the bottom of the scale, which means nothing was ever refused. The check was configured, documented, and doing nothing at all. Reading the config told you it was on. Only measuring told you it was off.
Not by intuition. The tool had been logging real queries, so the cutoff was derived from 103 captured questions across six real cohorts, hand-labelled into a 91-item set: 47 the manual can answer, 6 it cannot, and the rest keystroke fragments, template text and judgement calls that are excluded from scoring rather than quietly counted as wins.
Two things sit in front of an answer. A question check reads what is actually being asked. A support score rates how well the best passage in the text supports that exact question. The obvious assumption is that one of them is the real protection and the other is belt and braces. The measurement says otherwise.
A second model reads the question and a candidate passage together and rates how well that passage supports that exact question, on a scale from 0 to 1. Zero is unrelated. One means the passage answers precisely what was asked. It is not keyword overlap, and it is not the answer's confidence in itself. It is a judgement about evidence.
The cutoff sits at 0.20, and it can exist only because the two populations separate. Of the 47 questions the text can answer, 34 score above 0.8 and the median is 0.953. Questions the text cannot answer sit near the bottom. Without that separation no cutoff would be defensible, and a system with no cutoff always returns its best guess however poor it is.
| Question the text cannot answer | Support score | What stopped it |
|---|---|---|
| Is Social anxiety curable? | 0.3450 | question check |
| Is depression curable? | 0.0989 | question check |
| four spellings of a term absent from the text | 0.006 to 0.142 | support score |
The question check alone catches two of six. The support score alone catches five of six. Together they catch six.
Look at the top row to see why. "Is social anxiety curable?" scores 0.3450, comfortably above any cutoff that does not also start refusing legitimate questions. The manual describes a disorder in detail; it does not discuss cures. Relevance is high and answerability is zero, and only reading the intent separates them.
Now the bottom row. Those four are plain definition requests for a word the text does not contain. The question check has no reason to object to a plain definition request, so only the support score stops them.
Run either guard without the other and one of those questions reaches a student. That is not a tuning detail. Relevance and answerability are different properties, and a single confidence score cannot represent both.
Four answerable questions score below the cutoff. Three of them
(desl, alch, "Pulse of depression") were already
recorded as misses by the system before any threshold existed, so the cutoff is
not what loses them.
The fourth is real, and it is the one worth putting in front of you:
"kid has meltdowns way worse than normal tantrums" scores 0.0975 and is refused.
The manual covers this. It is disruptive mood dysregulation disorder. A student asked it in ordinary language, which is exactly the audience this tool exists to serve, and the cutoff turns them away.
That is the honest cost of the current setting. Refusal is not free: every threshold that stops a bad answer also stops some good ones, and the only useful question is which trade you are making and whether you measured it. Here the trade is one everyday-language question in forty-seven, and it is written down rather than discovered later by a user.
| Asked in ordinary words | Score | Resolved to |
|---|---|---|
| how do I tell bipolar 2 from MDD | 0.944 | p. 157 |
| difference between Bipolar 1 and 2 | 0.961 | pp. 155-156 |
| agoraphobia | 0.987 | p. 249 |
| self-esteem | 0.748 | p. 764 |
The obvious next step is to put a language model over the top, and it is worth being precise about what that buys and what it costs, because the usual claim is that it adds reasoning.
It adds synthesis. Comparing two disorders across four pages, restating criteria in plain language, working through an example. Those are real gains and students ask for exactly them.
What it does not do is remove the failure mode. It moves it. Retrieval alone fails visibly: you get a passage or you get a refusal. A language model over the same retrieval fails fluently, and a fluent wrong answer is the one that gets relied on. So the gate and the floor do not become less important once a model is added. They become the only thing standing between a good corpus and a confident mistake.
Which is the argument for the architecture rather than the model. The model is the part you can swap. The boundary, the citation, and the measured refusal are the parts that decide whether the answer can be trusted, and they are the parts that stay yours.
It is not a claim that hallucination has been eliminated. Nothing here measures that, and any vendor who says it has is telling you something they cannot support. What was measured is narrower and checkable: on this corpus with these 53 scored questions, the system refused every question it could not support and wrongly blocked none of the ones it could.
It is one corpus, one edition, one cohort of users, and a small sample. Fifty three scored questions means a single item moves the result by about two points, so the numbers are a floor on the evidence, not a precision claim. It has not been reproduced independently. And a single-source clinical manual sidesteps the jurisdiction and good-law problems that make legal retrieval harder, which is said here rather than left for you to notice.
Use the decision pack to compare vendors and discuss the next step with your decision-maker.
Get the decision pack · Request an assessment
← All resources · Back to Nucybersec · Privacy & measurement