The correct passage is never found, the reason is invisible unless you go and look at how the documents were chunked before searching.
"It says it doesn't know, but the answer is right there. I can see it in the document."
The system has 63 articles about the Solar System, and one of them contains this sentence, in full:
The three main rings are the narrow Adams Ring, 63,000 km from the centre of Neptune, the Le Verrier Ring, at 53,000 km, and the broader, fainter Galle Ring, at 42,000 km.
So it was asked:
How far is the Galle Ring from the centre of Neptune?
NOT IN CONTEXT
It can't find the answer, even though it's clearly in the corpus.
Searching doesn't happen over whole articles. An article is too big to hand to a model, and mostly irrelevant to any one question, so it's first chunked into short chunks of 900 characters each. Where the cuts land is a setting somebody chose.
Switch between the two settings below and watch what happens to the sentence.
Below is an extract from the article. The colour displays the 900-character chunks, and the red bar is where a chunk stops.
ncy band shows that it is a source of both continuous emission and irregular bursts. Both sources are thought to originate from its rotating magnetic field. In the infrared part of the spectrum, Neptune's storms appear bright against the cooler background, allowing the size and shape of these features to be readily tracked. == Satellite system and resonance == === Planetary rings === Neptune has a planetary ring system, though one much less substantial than that of Saturn and Uranus. The rings may consist of ice particles coated with silicates or carbon-based material, which most likely gives them a reddish hue. The three main rings are the narrow Adams Ring, 63,000 km from the centre of Neptune, the Le Verrier Ring, at 53,000 km, and the broader, fainter Galle Ring, at 42,000 km. A faint outward extension to the Le Verrier Ring has been named Lassell; it is bounded at its outer edge by the Arago Ring at 57,000 km. The first of these planetary rings was detected in 1968 by a team led by Edward Guinan. In the early 1980s, analysis of this data along with newer observations led to the hypothesis that this ring might be incomplete. Evidence that the rings might have gaps first arose during a stellar occultation in 1984 when the rings obscured a star on immersion but not on emersion. Images from Voyager 2 in 1989 settled the issue by showing several faint rings. The outermost ring, Adams, con
No chunk anywhere contains both “Galle Ring” and “42,000 km”.
The cut has landed between the two words of the name. One chunk stops after “Galle”; the next opens with “Ring, at 42,000 km” and gives a distance for a ring it can't name.
Asked “How far is the Galle Ring from the centre of Neptune?”, with the 4 best-matching chunks handed to it, the model said:
Before it's searched, a document is chunked. Chunk overlap lets each chunk repeat a bit of the chunk before it, so something landing on a cut still appears whole somewhere. With no overlap the cut fell between the two words of the ring's name, and the phrase “Galle Ring” existed nowhere in the system.
It isn't free. Repeating text means more chunks to store and search, 2478 became 2948, and those near-duplicates compete with each other in the results. One of the five test questions on these same articles fell from position 2 to outside the top ten altogether as a direct result of switching overlap on.
Or cut where the document already breaks. Work out its own boundaries first and cut on those, so a paragraph stays a paragraph and a table stays a table.
This costs a parsing pass. On a document system I work on, built with RAGFlow, the structure-aware parser reads layout and tables before anything is indexed. These bigger chunks change what's retrieved, so it's a switch that requires re-measurement rather than an assumption. Understanding how the business works and the problems it needs to solve helps to accurately improve the system. However, for the number of documents we needed to process, we needed more power. Reading layout and tables is model work, so it required a GPU rather than a spare thread. The live box runs five 24 GB RTX 3090s, two Xeon Gold 6248R processors (48 cores, 96 threads), 384 GB of RAM and a 2 TB NVMe drive.
This entry is reasoning about how other tools work, not something measured here. Everything else on this page came out of running it.
No. Cuts moving with the overlap belongs to splitters that take a size and an overlap - the default nearly everywhere, which is why it matters - and not to ones that cut at sentence or paragraph ends, or follow a document's own headings, or offer no overlap setting at all. A splitter that cuts where the writing already pauses is much less likely to cut through a name to begin with. To find out which kind you have: chunk one document at two different overlap settings and see whether the second cut moved. If it did, yours works like this one. There's also a different technique worth knowing, sometimes called context expansion or parent-document retrieval, which keeps the cuts where they're and instead hands the model a wider window around whichever chunk matched - arguably the cleaner fix for this exact failure.glm-5.1:cloud, 3 samples per setting.