Results

Everything here is read from the saved evaluation runs, not typed in. Stages that call a language model were repeated and are shown as mean ± standard deviation, so you can see which differences are real and which are noise.

Retrieval quality by stage

Full-hit@5 is the share of questions where every passage needed to answer was in the top 5. Bars show the mean; the whisker is one standard deviation across runs.

11 IETF RFCs, 20 hand-labelled questions, 5 runs · shipped defaults

Loading…
StageRecall@5Full-hit@5MRRRetrieval p50p95

* These rows include neighbour expansion (each hit is widened by the chunk before and after it), so they are measured over more text than the rows above them; compare them by answer accuracy below. Lexical search is the strongest single retriever on the synthetic set but the weakest on real text, and absolute quality is much lower on real documents. That gap is why the synthetic numbers alone would have been misleading.

Answer quality and the agent loop

Answers are graded by a language model against a gold answer. The judge is the same model family as the answerer, so treat accuracy as secondary to the retrieval numbers above.

DatasetPipelineAnswer accuracyUnsupported claimsMulti-part full-hitEnd-to-end p50p95
Real documentspipeline, no agent–––––
Real documentspipeline + agent loop–––––
Synthetic corpuspipeline, no agent–––––
Synthetic corpuspipeline + agent loop–––––

The agent loop clearly helps questions with several parts. On the synthetic set, where questions are easy, it is not distinguishable from the plain pipeline and costs roughly three times the tail latency; the benefit shows up on harder, real text.

Cost per question

Token counts come from the model provider for every eval run. Dollar figures use the price table in configuration and appear only for models with a price set.

No token usage recorded yet for this dataset.

How far to trust this

  • The real-document set is 20 questions; one question is 5 points of full-hit, so differences of a few points are noise.
  • Those questions were written by a language model from verbatim RFC sentences, and the judge is the same model family, not an independent human.
  • A retrieved passage counts as a hit only if it contains the labelled quote, so a different passage that also answers the question counts as a miss. Retrieval numbers are a lower bound.
  • Latency was measured with 4 concurrent queries against a remote database, which inflates it.