$ retrieval-augmented generation

RAG, measured against ground truth

Ask a question with and without a corpus. The model never changes — only what it is allowed to read. Where the corpus ships a known-correct answer, every question is scored: you can see exactly which ones the retriever gets right and which it misses.

Ask

Retrieved passages

Your own questions, over a baseline of runs from the author. Nobody sees anyone else's.

The same retriever, evaluated offline against all 10,570 labelled SQuAD questions and both encoders — 21,140 scored retrievals. The analysis tab above shows what this demo has been asked; this one shows how the retriever behaves across the whole benchmark.