Cut in half by a chunk boundary

The correct passage is never found, the reason is invisible unless you go and look at how the documents were chunked before searching.

"It says it doesn't know, but the answer is right there. I can see it in the document."

A real example

The system has 63 articles about the Solar System, and one of them contains this sentence, in full:

The three main rings are the narrow Adams Ring, 63,000 km from the centre of Neptune, the Le Verrier Ring, at 53,000 km, and the broader, fainter Galle Ring, at 42,000 km.

So it was asked:

How far is the Galle Ring from the centre of Neptune?

NOT IN CONTEXT

It can't find the answer, even though it's clearly in the corpus.

What actually happened

Searching doesn't happen over whole articles. An article is too big to hand to a model, and mostly irrelevant to any one question, so it's first chunked into short chunks of 900 characters each. Where the cuts land is a setting somebody chose.

Switch between the two settings below and watch what happens to the sentence.

Below is an extract from the article. The colour displays the 900-character chunks, and the red bar is where a chunk stops.

ncy band shows that it is a source of both continuous emission and irregular bursts. Both sources are thought to originate from its rotating magnetic field. In the infrared part of the spectrum, Neptune's storms appear bright against the cooler background, allowing the size and shape of these features to be readily tracked. == Satellite system and resonance == === Planetary rings === Neptune has a planetary ring system, though one much less substantial than that of Saturn and Uranus. The rings may consist of ice particles coated with silicates or carbon-based material, which most likely gives them a reddish hue. The three main rings are the narrow Adams Ring, 63,000 km from the centre of Neptune, the Le Verrier Ring, at 53,000 km, and the broader, fainter Galle Ring, at 42,000 km. A faint outward extension to the Le Verrier Ring has been named Lassell; it is bounded at its outer edge by the Arago Ring at 57,000 km. The first of these planetary rings was detected in 1968 by a team led by Edward Guinan. In the early 1980s, analysis of this data along with newer observations led to the hypothesis that this ring might be incomplete. Evidence that the rings might have gaps first arose during a stellar occultation in 1984 when the rings obscured a star on immersion but not on emersion. Images from Voyager 2 in 1989 settled the issue by showing several faint rings. The outermost ring, Adams, con

No chunk anywhere contains both “Galle Ring” and “42,000 km”.

The cut has landed between the two words of the name. One chunk stops after “Galle”; the next opens with “Ring, at 42,000 km” and gives a distance for a ring it can't name.

Asked “How far is the Galle Ring from the centre of Neptune?”, with the 4 best-matching chunks handed to it, the model said:

  • NOT IN CONTEXT
  • NOT IN CONTEXT
  • NOT IN CONTEXT
Full written explanation
Symptom
"It says it doesn't know, but the answer is right there. I can see it in the document."
What's happening
Documents are chunked before being searched. A chunk boundary landed in the middle of the fact being asked about, so the name of the thing ended up in one chunk and its value in the next. Neither chunk answers the question, and no amount of searching can return a chunk that doesn't exist.
Why it's hard to spot
Every check you would normally run says the system is healthy. The document is indexed. The search returns results quickly, and they're about Neptune's rings - they look right. The model behaves impeccably: handed passages that don't contain the answer, it declines to invent one and says so. Nothing anywhere reports an error. The only place the failure is visible is in the chunks themselves, which is the one thing most systems never show you.
The check
Search your own index for the exact phrase that answers the question. If it isn't there - not ranked low, but absent - the answer was destroyed at the chunking stage, before search was ever involved.
The fix, and its cost
Chunk overlap, set longer than the longest fact you need kept whole - a fact was destroyed only when the gap between its name and its value exceeded the overlap, at every setting measured here. It costs storage and index size (2478 chunks become 2948 at 150 characters of overlap), and it costs precision, because near-duplicate chunks crowd each other in the rankings. Measured here: one of five test questions dropped out of the top three results as a direct result of adding overlap. Splitting on sentence or paragraph boundaries instead of a fixed character count should avoid the problem more cleanly, at the cost of chunks of uneven size - though that one is reasoning rather than something measured here, and it's on the list to test.
Why more results don't help
The usual first move is to ask for more results back. It can't work here. With no overlap the answer is in none of the 2478 chunks, so there's nothing for a longer list to reach - checked at 1, 3, 5, 10 results, missed at every one. That's worth separating from the usual case, where more results do help because the answer was merely ranked too low.
How much of this is luck
A fair question, and the answer is: with no overlap, most of it. Six overlap settings were measured, and the answer only lands inside the repeated stretch at the largest of them - the one shown above. At 25, 50, 100 and 150 characters it survives for a different reason: raising the overlap also moves every cut in the document, so the cut simply lands elsewhere. Each chunk is still 900 characters, but it now starts earlier, inside the chunk before it, so it advances by 900 minus the overlap and every cut after the first shifts. On this article the first cut sits at character 899 whatever the setting, while the second moves from 1,799 with no overlap to 1,500 at 300.
So what is the evidence
Not one sentence - the whole corpus. 943 facts of this shape appear across the five articles: a distinctive name and a value stated together. Chunking destroys 71 of them with no overlap and 0 at 300 characters, and not the same ones each time. The rule behind it held at every setting measured: a fact is destroyed only when the gap between its name and its value is longer than the overlap. “Galle Ring” to “42,000 km” is 24 characters, so any overlap above that makes this answer impossible to lose. The practical version isn't "turn on overlap" but: find the longest fact you need kept whole, and set the overlap above it.
Does every chunker behave this way

This entry is reasoning about how other tools work, not something measured here. Everything else on this page came out of running it.

No. Cuts moving with the overlap belongs to splitters that take a size and an overlap - the default nearly everywhere, which is why it matters - and not to ones that cut at sentence or paragraph ends, or follow a document's own headings, or offer no overlap setting at all. A splitter that cuts where the writing already pauses is much less likely to cut through a name to begin with. To find out which kind you have: chunk one document at two different overlap settings and see whether the second cut moved. If it did, yours works like this one. There's also a different technique worth knowing, sometimes called context expansion or parent-document retrieval, which keeps the cuts where they're and instead hands the model a wider window around whichever chunk matched - arguably the cleaner fix for this exact failure.
What doesn't work
Asking for more results back, a better embedding model, or better prompting. All three assume the answer is somewhere in the index and merely undervalued. It isn't there at all.
How this was measured
Retrieval only, over the pinned corpus, with the same embedding model the rest of the pipeline uses. A chunk counts as answering the question only if it contains both the name of the thing and the value being asked for: a chunk saying something is 42,000 km away, without saying what, does not answer 'how far is the Galle Ring'. Chunk size is held at 900 characters throughout; overlap is the only thing that varies. The question was found by searching the pinned corpus for any fact whose name and value fall in different chunks at zero overlap - not chosen to make the point. Answers generated by glm-5.1:cloud, 3 samples per setting.