Two sources disagreed, so it invented a reason why

"It gave me two different numbers in the same sentence and made it sound like both were right. I only noticed because I already knew the answer."

Two documents, both sure of themselves

The system has two documents that answer the same question:

"Jupiter has 95 moons that are officially recognized by the International Astronomical Union."

Jupiter Facts - NASA Science · last updated April 2025 · source

"Jupiter has 115 known natural satellites, and it is likely that this number will go up due to increasing telescopic observations."

At the time which these documents were created, they were correct. But, the number of discovered satellites increased between writing

How did it answer

How many moons does Jupiter have?

Jupiter has 95 officially recognized moons, though at least 115 moons orbit the planet.

Both numbers count the same moons. One document isn't more officially recognised than the other.

Try the fixes

Of the 36 combinations of chunk size, overlap and number of results measured, 32 return both figures, so this isn't a corner case. Switch between the three columns and watch the top number.

7
of 32 answers invented a distinction between the two figures and stated it as fact
Said the sources disagree14
Gave one figure, said nothing about the other11
Invented a distinction7

Retrieval is identical in all three. Same passages, same order, same scores. Nothing a retrieval score can measure moves at all between these.

What it actually answered

Five of the 32, worst first. Every answer at every setting is in the results file.

Why no retrieval score would have caught this

It's sometimes fixed with the date alone, as the LLM spots the difference between the passages and intelligently choses the most recent one. But, it's greatly improved with a strict instruction to do so.

Hit rate, rank and precision stay perfect. The right passages came back. The problem is what the system said about them, and those scores never look at the answer.

An evaluation that checks the answer catches this. It needs someone to record what a correct answer has to contain. Such evaluation led to changes in my own work. I work on a retrieval system which covers engine manuals. They get reissued, and each one covers multiple engine types, sizes etc. The parse now enriches each document with the engine types it covers, and the LLM is prompted to check those against the engine the question is about. It's similar to the date issue presented in this case, we enrich the document with the date and prompt the LLM to use the most recent.

The full diagnosis
Symptom
An answer contains two different figures for one thing, joined by a confident explanation of why they differ, and the explanation is invented.
What's happening
Two documents in the collection disagree, because one was updated and the other was not. Both are retrieved. Neither carries anything saying which is current, so the system reconciles them itself, and the most fluent way to reconcile two numbers is to assume they measure different things.
Why it's hard to spot
Retrieval succeeded, so every retrieval metric is green. The answer isn't a refusal and not obviously absurd. It's more detailed than a correct answer would be, which makes it read as better rather than worse.
The check
When an answer explains why two figures differ, go and read both passages before you believe the explanation. If it came from the documents you will find it there. If it didn't, it was invented in the last sentence.
The fix, and what it costs
Attach the last-updated date to each passage after it has been retrieved, and tell the system to say when passages disagree and prefer the most recent. The timing matters: do it before the documents are indexed and you change what gets found instead. Both halves of the fix are needed, and they do different jobs. The dates alone cut the invented distinctions from 7 to 2 of 32, which is most of the danger gone, but they barely change how often the system mentions the disagreement at all: 14 of 32 without them, 17 with. Adding the instruction takes invented distinctions to 0 and gets all 32 to say the sources disagree. The cost is that you need reliable dates, which plenty of document stores don't have, and some of your instruction budget.
What doesn't work
Asking for more results. Wherever both figures already come back, going deeper only adds more of them, and both figures already come back at 32 of the 36 settings measured. This isn't a case where the right passage was missing.
The honest limit of this case
The two sentences do use slightly different words, "officially recognized" against "known", and that difference is the hook the invented distinction hangs on. That's not a flaw in the example. Real sources phrase things differently, and that's exactly what gives a system room to paper over a conflict.
How this was measured
3 chunk sizes against 3 overlaps against 4 result counts, 36 combinations, each asked three ways: as retrieved, with dates attached, and with dates plus an instruction. Every answer is kept word for word in the results file so the counts above can be checked rather than trusted. Both documents are the real published ones, unedited.