This one is the search itself, and it's the only failure which the language model plays no part in.
"I typed the reference number that's printed in our document. It came back with something else entirely."
The code 2008 LC18 appears in these documents. Here's what the search returned asked for:
What is 2008 LC18?
NOT IN CONTEXT
The passage holding it came 139th out of 2,948.
Ask for the same code again, but provide a clue about the code:
What is 2008 LC18?
139th
What is the Neptune trojan 2008 LC18?
1st
It can find the code once you tell it what the code is. However, the problem is that if you could describe the thing, you wouldn't be looking up its number.
Semantic search works by meaning, not by matching letters. It turns your question into a position in a space of meanings and finds passages that sit nearby. 2008 LC18 means nothing, so it lands nowhere in particular.
Plain keyword search finds every one of the 4 codes at or near the top. The keyword search algorithm used here is BM25, the standard keyword ranking behind most search engines and most RAG systems. Then you ask it something in your own words:
What is the enormous swirling weather system on Jupiter called?
the passage that answers it came 872nd
Keyword search matches the words you typed. Ask a question that happens to use none of the words the document used, and it has nothing to go on. Each method fails in it's own way.
All 7 searched against the same 2,948 passages: 4 exact codes, and 3 questions asked in your own words. Two controls: how much of the vote each search gets, and how many results are kept and passed on. One combination answers all 7.
3 of 7 questions get the right passage
You need both. Without keyword the codes are lost, without semantic the plainly worded questions are. But averaging an excellent ranking with a useless one has a cost. The jupiter storm question sits 1st by semantic and 872nd by keyword, so the blend puts it 23rd. Better than 872nd. Still off a short list.
Tuning the balance won't fix it. Below 25 results, every setting of the first slider returns the same 4 of 7. Pool size is the lever, not the blend.
That's why real systems keep more results than they show you. Retrieve a wide pool, then re-sort it with a slower model that reads each passage properly. The blend gets the right passage into the pool. Ordering it is a separate job, this is known as re-ranking.
Or don't blend at all. Send anything code-shaped to keyword, everything else to semantic. That gets all 7 inside the top ten without widening anything. The price is a rule to maintain, and a query that's half code and half question falls between the two.
| Question | Sent to | Routed | Even blend |
|---|---|---|---|
| What is EETA79001? | keyword | 4th | 2nd |
| What is HD 209458 b? | keyword | 1st | 1st |
| What is 136199? | keyword | 1st | 1st |
| What is 2008 LC18? | keyword | 1st | 4th |
| How long does it take Mars to travel once around the Sun? | semantic | 1st | 17th |
| Why did Pluto stop counting as a planet? | semantic | 1st | 12th |
| What is the enormous swirling weather system on Jupiter called? | semantic | 7th | 23rd |
Which approach is right depends on your documents and your users. This corpus is 4 codes against 3 ordinary questions, which is what makes routing look this good here. The only way to know yours is to measure your RAG pipeline and its corpus on the questions people actually ask, use evaluation.
The two searches are built on opposite assumptions, and each one's strength is exactly the other's weakness.
Semantic search
Good: finds the answer when you phrase the question your own way.
Bad: a code has no meaning, so there's nothing for it to be near.
Keyword search
Good: a rare string of characters is the easiest thing in the world to match exactly.
Bad: change the words and it has nothing to match.
Measured over 2,948 passages, chunked at 900 characters with 150 overlap. Keyword ranking is BM25Okapi; the two rankings are combined with reciprocal rank fusion, constant 60. The 3 questions asked in your own words are facts this model already knows, so for those the evidence is where the passage ranked, not whether the answer came back right.