Results
Everything here is read from the saved evaluation runs, not typed in. Stages that call a language model were repeated and are shown as mean ± standard deviation, so you can see which differences are real and which are noise.
Retrieval quality by stage
Full-hit@5 is the share of questions where every passage needed to answer was in the top 5. Bars show the mean; the whisker is one standard deviation across runs.
11 IETF RFCs, 20 hand-labelled questions, 5 runs · shipped defaults
| Stage | Recall@5 | Full-hit@5 | MRR | Retrieval p50 | p95 |
|---|
* These rows include neighbour expansion (each hit is widened by the chunk before and after it), so they are measured over more text than the rows above them; compare them by answer accuracy below. Lexical search is the strongest single retriever on the synthetic set but the weakest on real text, and absolute quality is much lower on real documents. That gap is why the synthetic numbers alone would have been misleading.
Answer quality and the agent loop
Answers are graded by a language model against a gold answer. The judge is the same model family as the answerer, so treat accuracy as secondary to the retrieval numbers above.
| Dataset | Pipeline | Answer accuracy | Unsupported claims | Multi-part full-hit | End-to-end p50 | p95 |
|---|---|---|---|---|---|---|
| Real documents | pipeline, no agent | – | – | – | – | – |
| Real documents | pipeline + agent loop | – | – | – | – | – |
| Synthetic corpus | pipeline, no agent | – | – | – | – | – |
| Synthetic corpus | pipeline + agent loop | – | – | – | – | – |
The agent loop clearly helps questions with several parts. On the synthetic set, where questions are easy, it is not distinguishable from the plain pipeline and costs roughly three times the tail latency; the benefit shows up on harder, real text.
Cost per question
Token counts come from the model provider for every eval run. Dollar figures use the price table in configuration and appear only for models with a price set.
How far to trust this
- The real-document set is 20 questions; one question is 5 points of full-hit, so differences of a few points are noise.
- Those questions were written by a language model from verbatim RFC sentences, and the judge is the same model family, not an independent human.
- A retrieved passage counts as a hit only if it contains the labelled quote, so a different passage that also answers the question counts as a miss. Retrieval numbers are a lower bound.
- Latency was measured with 4 concurrent queries against a remote database, which inflates it.