The explanation that never gets retrieved

A large document explains how to read its tables near the front. The tables themselves are dozens of pages later. Both parts get indexed. The LLM can't read the table as it isn't aware of the table instructions.

A real example

We have an inspection schedule in a table. Every row in this inspection schedule gives a part, an interval, and a one-letter code. The code says what actually triggers the inspection. Here's the row the search returned, first out of 2,967:

Subsystem Inspection Schedule
Subsystem: Attitude thruster valve seals
Interval: 4,000 operating hours
Basis: R

What can the LLM interpret from this table row?

For the attitude thruster valve seal inspection, what does a Basis of R mean?

NOT IN CONTEXT

The letter R is explained: It's explained in the introduction of the same document, in a legend which explains R. That section is indexed too. It came back 13th.

Why the table instructions never arrive

With the instructions on how to read table data within the introduction and the table in the middle of the document, the LLM can't read understand the table rows when they are retrieved. Search returns a ranked list, not a set of related things. So the way to solve this is to enrich each table found in the document with the legend found in the introduction.

Try it

Two controls. How many results you keep, and whether the row carries its own explanation. Watch the search score on the left while you change them.

What the search scored

1st

the row came back at rank 1 of 2,967

Unchanged by either control.

What came back

NOT IN CONTEXT

The introduction is rank 13, so it didn't come back.

The table row the search returned

Subsystem Inspection Schedule
Subsystem: Attitude thruster valve seals
Interval: 4,000 operating hours
Basis: R

What this one doesn't do

It refuses rather than guessing what R means, which is the safe way to fail. I took the refusal instruction out and it continued to decline.

Worth contrasting with the case next door, where the same model cheerfully supplied a number of its own rather than admit it hadn't found one. The difference is whether it has an opinion to fall back on. Nobody has a prior belief about what R means in someone's maintenance schedule.

Full written explanation
Symptom
"Our search scores are perfect. It still says it can't answer."
What's happening
A document explains how to read its tables in one place and prints the tables in another. Chunking separates them. Retrieval returns a ranked list rather than a set of related pieces, so nothing guarantees the explanation travels with the row. Because the explanation is written to cover every row, it shares almost no wording with a question about one of them, and never ranks high enough to be returned.
The check
Look at where the explaining passage actually ranks for the question you care about, not whether it's in the index. If the row is first and the explanation is thirteenth, your cut-off decides whether the system can answer, and your retrieval scores will look identical either way.
The fix, and its cost
At build time, copy the sentences that explain a table into each row of it, filtered to the codes that row actually uses. Costs a preparation step, some duplication across chunks, and a little room in each one. In exchange the row means something wherever it is sent, and nothing downstream has to be lucky.
What doesn't work
Putting the explanation in its own tidy chunk, which still has to win a ranking it can't win. A better embedding model, which is ranking the same mismatched text. Raising the cut-off, which works only while the collection is small enough for the explanation to be within reach.

The schedule, the part names, the codes and the figures are all invented, and no material from the real system this is drawn from appears here. The invented handbook is 19 passages, searched alongside 2,948 real ones from this site's corpus so the ranking is measured against genuine competition. Answered by glm-5.2:cloud.