A long pile of text buries an answer, as seen in Lost in the middle.
"We spot-checked one of the figures it gave us. It was wrong."
This system was asked for a number. The document holding that number was found by the search and handed to the model, in a pile of 850 passages. Here's what came back, on 9 of 10 attempts:
What is the density of Uranus's moon Oberon, in grams per cubic centimetre?
1.63 g/cm³
That number appears in none of the documents. It's what this model says when asked the question and give it nothing at all to read. RAG is designed to overcome hallucinations, but if it falls back to the model this issue will remain. The real figure, the one in the document, is 1.68, it almost looks correct.
When this system fails to read the passage it was given, it does one of two things, and which one depends entirely on whether the model had an opinion of its own to fall back on.
A code the model has never seen
60 of 60
failures said I can't find it. Nothing to fall back on, so it owned up.
A figure it misremembers
20 of 21
failures produced a number anyway. It had its own answer ready, so it provides you that.
Same pile, same positions, same model. The only difference is whether it already thought it knew.
The second one is the dangerous one. A refusal is easy to notice while a wrong number is silent. It can produce confident answers in the correct format, within a plausible range. Nothing in the answer marks it out, and the search logs are clean because the search worked.
Below are four questions asked of the same pile. Two ask for an ordinary sentence and two for a bare number or code. Two ask about something invented for this experiment, which the model can't have seen before, and two about something real, which it may already think it knows.
Drag the slider to make the pile bigger, and watch all four at once. Two of them fall apart. The interesting part is which two, and what each one does on its way down.
about 2,181 of 202,752 tokens the model can hold - 1.1% of the window
Each bar shows the worst spot in the pile for that particular thing, because the two that fail don't fail in the same place. The full set of positions is in the table below.
What is the calibration code for the Vega-7 probe?
What was the Kestrel-3 mission designed to study?
What is the density of Uranus's moon Oberon, in grams per cubic centimetre?
How long does it take Mars to orbit the Sun?
The comparison above shows each thing at its worst position. Here is the same pile with the answer at every position measured, so nothing is being chosen to flatter the point. Each figure is how many of the attempts found the answer. Note where the failures sit: the code is lost in the middle, the misremembered figure at the start.
| Looking for | 0% | 22% | 44% | 78% | 100% |
|---|---|---|---|---|---|
| A code the model has never seen | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 |
| A sentence the model has never seen | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 |
| A figure the model misremembers | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 |
| A fact the model knows well | 10/10 | 10/10 | 10/10 | 10/10 | 9/10 |
A sentence is supported by it's collection of words, so half-reading it still leaves enough to work with. However for a number or identifier if one digit is wrong, there's nothing left to reconstruct it from.
An ordinary sentence, partly lost
Its purpose was to study seasonal storms.
You can still fill that in. So can the model.
A code, partly lost
XQ-47
Nothing to fill it in from. It's simply gone.
The model has been previously trained. Faced with a code it hasn't been trained on, it said so. Faced with a figure it thought it already knew, it can confidently provide inaccurate data.
This is exactly what a lookup system is built for. Order numbers, case references, part codes, dates: the things people actually go to a document store to retrieve are precisely the kind of content most likely to be dropped from a long context, and dropped fastest when surrounded by similar content.
Related case: A code semantic search can't find is the same fragility at the opposite end of the pipeline. Semantic search can miss the code all together...