The search does its job perfectly. The answer is wrong anyway.
"It retrieves the right document and still gets the answer wrong."
The system was prompted and the golden chunk, the one I'd recorded in advance as holding the answer, was retrieved. Here's the result:
What is the calibration code for the Vega-7 probe?
NOT IN CONTEXT
This was staged on purpose so that the 'lost in the middle' issue can be fully illustrated. The answer passage was moved through the pile, first to last, with everything else remaining in place, it's the only way to show the effects of the passage moving between the start and the end, while displaying the effects of the middle.
Below are two sliders: how many passages the model was given, and where the answer sat in the pile. Drag either one. Every combination is real, with the actual passage and answer behind it.
71,402 of 202,752 tokens the model can hold - 35.2% of the window
100% at the start, 100% at the end, and as low as 20% at 67% through.
Found in 2 of 10 attempts (20%)
The Vega-7 orbiter was inserted into a polar survey orbit following a single mid-course correction, and its instrument suite was brought online over the following eleven days. Ground controllers assign every survey platform a calibration code used to reconcile telemetry across mission phases, and the calibration code for the Vega-7 probe is XQ-4417. Subsequent passes returned atmospheric density profiles consistent with earlier modelling, though the upper-altitude readings showed a seasonal variation that had not been predicted. The survey continued for a further two years before the platform was placed into a reduced-power mode, with periodic contacts maintained from the deep space network throughout.
Two fixes that don't touch retrieval: move the answer nearer the start or end of what the model is given, which held at every pile size measured here, or cut the pile right down: at 10 passages nothing was missed anywhere in it. Both have tradeoffs: the first isn't always controllable in a live system; and the second means leaving material out.
Language models pay closest attention to the start and the end of what they're given, and least to the middle. This is called lost in the middle. Nothing about search caused it. The passage was found; it just wasn't read carefully once it got there.
No retrieval metric shows this. Recall, precision, ranking - all can be perfect while the answer is wrong. That's why it goes unnoticed in real systems.
More context isn't free. Adding passages "to be safe" can bury the one that mattered. This runs against most people's instinct that more information can only help.
It's a probability, not a switch. In this data the middle of a 320-passage pile found the answer 46% of the time, and the worst position in it managed 20% across 10 attempts.
The ends are the reliable part. Across every pile size measured here, up to 850 passages filling 88% of everything this model can hold, the first and last positions found the answer every single time. The middle didn't.
What doesn't work: retrieving more passages; reranking; or a better embedding model. The issue will remain if these are tried.