Every case on this site hands you a system that's giving the wrong answer and asks you to work out why. This page explains what the system is actually doing, and how you would know whether it.
Under the hood
There are 63 documents - Wikipedia articles about the Solar System, from the Sun and the planets to the moons, the dwarf planets and the missions sent out to look at them. That's everything the system knows. It has never seen anything else, and there's far too much of it to hand over all at once, which is why it has to search.
Each document goes through a RAG pipeline. Each case has a slight variation of how RAG can be implemented. I won't got into detail on how a full RAG pipeline works as there's plenty of information about it out there already.
For each question, I already know which chunk holds the real answer, because I read the documents and recorded it. For each case we can then check to see if the correct chunk (know as the golden chunk) has been retrieved.
A brief example case
Alter the characters and overlap
80%
of questions answered by a chunk that genuinely contains the answer · mean reciprocal rank 0.65 · precision at 4 25%
We can learn what happens to the likelihood of retrieving the correct chunk when changing the characters and overlap of our chunks. This specific example is covered in more detail within the actual cases.
How this was built
Documents
63
Model
sentence-transformers/all-MiniLM-L6-v2
Model version
1110a243fdf4
Built
09/09/2026, 14:50:14
Every number in this project was produced by running a real retrieval pipeline over the pinned corpus.
The model that writes the answers has changed once already. Cases built before 25 September 2026 were measured against glm-5.1:cloud, which was retired that day. Those numbers are a historical record and aren't being re-run: they still stood on real measurement when they were made, and a model's retirement doesn't undo that. Cases built after that date use glm-5.2:cloud, which has a context window over five times larger. This has been noted for each case that this applies to.
What was searched
63 Wikipedia articles about the Solar System: the Sun and the planets, the dwarf planets, the major moons, the belts and clouds between them, as well as some of the missions. Together they come to 2.2 million characters, which is far more than fits in the model's context window. Someone couldn't simply upload all of these files to an LLM as they wouldn't fit.
These are the copies, taken on a fixed date and stored with this site. The corpus is the exact files found here, not the live sites.
Text from Wikipedia, reused under CC BY-SA 4.0. Each "original" link above points at the exact revision used, which is what that licence asks for and what makes these results reproducible.