An answer pulled from memory, not the document

A long pile of text buries an answer, as seen in Lost in the middle.

"We spot-checked one of the figures it gave us. It was wrong."

A real example

This system was asked for a number. The document holding that number was found by the search and handed to the model, in a pile of 850 passages. Here's what came back, on 9 of 10 attempts:

What is the density of Uranus's moon Oberon, in grams per cubic centimetre?

1.63 g/cm³

That number appears in none of the documents. It's what this model says when asked the question and give it nothing at all to read. RAG is designed to overcome hallucinations, but if it falls back to the model this issue will remain. The real figure, the one in the document, is 1.68, it almost looks correct.

The part that decides whether you ever find out

When this system fails to read the passage it was given, it does one of two things, and which one depends entirely on whether the model had an opinion of its own to fall back on.

A code the model has never seen

60 of 60

failures said I can't find it. Nothing to fall back on, so it owned up.

A figure it misremembers

20 of 21

failures produced a number anyway. It had its own answer ready, so it provides you that.

Same pile, same positions, same model. The only difference is whether it already thought it knew.

The second one is the dangerous one. A refusal is easy to notice while a wrong number is silent. It can produce confident answers in the correct format, within a plausible range. Nothing in the answer marks it out, and the search logs are clean because the search worked.

Four things to look for, in the same pile of text

Below are four questions asked of the same pile. Two ask for an ordinary sentence and two for a bare number or code. Two ask about something invented for this experiment, which the model can't have seen before, and two about something real, which it may already think it knows.

Drag the slider to make the pile bigger, and watch all four at once. Two of them fall apart. The interesting part is which two, and what each one does on its way down.

about 2,181 of 202,752 tokens the model can hold - 1.1% of the window

0 tokens202,752 tokens (everything this model can hold)

Each bar shows the worst spot in the pile for that particular thing, because the two that fail don't fail in the same place. The full set of positions is in the table below.

  1. A code the model has never seen10 of 10

    What is the calibration code for the Vega-7 probe?

  2. A sentence the model has never seen10 of 10

    What was the Kestrel-3 mission designed to study?

  3. A figure the model misremembers10 of 10

    What is the density of Uranus's moon Oberon, in grams per cubic centimetre?

  4. A fact the model knows well9 of 10

    How long does it take Mars to orbit the Sun?

Every position measured at this pile size

The comparison above shows each thing at its worst position. Here is the same pile with the answer at every position measured, so nothing is being chosen to flatter the point. Each figure is how many of the attempts found the answer. Note where the failures sit: the code is lost in the middle, the misremembered figure at the start.

Looking for0%22%44%78%100%
A code the model has never seen10/1010/1010/1010/1010/10
A sentence the model has never seen10/1010/1010/1010/1010/10
A figure the model misremembers10/1010/1010/1010/1010/10
A fact the model knows well10/1010/1010/1010/109/10
Full written explanation
Symptom
"We spot-checked one of the figures it gave us. It was wrong."
What's happening
Two things, one after the other. As the pile of text grows, the model reads ordinary prose more reliably than bare values: a code or a figure has no surrounding meaning to hold it up, so it's the first thing lost. Then, having lost it, the model answers from what it already believed. On an invented code it has no belief to fall back on and says it can't find the answer. On anything real it usually does, so it fills the gap instead of reporting it.
The check
Ask the question with the documents removed. If the model produces a confident answer with nothing to read, it has an opinion of its own, and any answer it gives you with documents may be that opinion rather than your data. Then ask for an ordinary fact and an exact value from the same document at the same pile size: if the fact survives and the value doesn't, the content type is the variable, not the search.
The fix, and its cost
Shrink the pile, keep exact values near the edges of what is assembled, or attach a value to a sentence that explains it (the enrichment approach used for the mangled-table case). Each costs either completeness or extra authoring work. None of them tells you when it has gone wrong, which is why the check above matters more than any of them.
What doesn't work
A better embedding model or more insistent prompting. The search already found the passage and handed it over; the loss happens afterwards. Asking the model to say when it's unsure doesn't help either: on these failures it wasn't unsure, it was wrong.