Architecture
Fathom is a retrieval-augmented chat system: it finds the passages in your documents that bear on a question, reasons about whether it has enough, and writes an answer that cites where each claim came from. Every stage is measured, so the trade-offs below have numbers behind them.
What happens to a question
Green stages run in Postgres, purple stages call a language model. The agent loop (plan, reflect, follow-up) can be switched off in the sidebar.
- 1QuestionUp to 1,000 characters, with the last 6 chat turns
- 2PlanMulti-part questions are split into up to 4 standalone searches; simple ones skip this call
- 3Hybrid searchPostgres full-text + pgvector, in parallel per sub-question, fused with Reciprocal Rank Fusion
- 4RerankAn LLM scores the top 20 fused passages 0–9 in one call; the best 5 are kept
- 5ReflectIs every entity in the question covered? If not, a follow-up search (one by default, configurable)
- 6EvidenceDe-duplicated and capped at 8 passages, optionally widened with neighbouring chunks
- 7AnswerStreams a cited answer that may use only the numbered sources
Every request gets a trace id, stage timings, token counts and a dollar estimate, shown under each answer and written to the server log as one JSON line.
Ingestion and storage
- Reads .txt .md .pdf; the original is kept in S3 when configured.
- Splits along paragraph, then sentence boundaries into ~900-character chunks with 150 characters of overlap.
- Embeds chunks in batches with text-embedding-3-small (1536 dimensions).
- Scans the text for instruction-like content and flags the document (see Security).
- Optional contextual headers: each chunk is prefixed with its document description and nearest heading.
- documents: title, source, S3 key, ingest scan flags.
- chunks: content, a vector(1536) embedding, and a generated tsvector.
- HNSW index (cosine) for vectors, GIN index for full-text, cascade delete from document to chunks.
- Both retrieval paths and the metadata live in one place, so a delete is one statement and there is no index to keep in sync.
Retrieval
Three ideas, each kept because it measurably helped: search two ways, rerank a wide pool, and let an agent loop handle questions with several parts.
Full-text catches exact names and numbers; vectors catch paraphrase. Their ranked lists (20 each) are merged with Reciprocal Rank Fusion, which needs no score calibration.
One call scores the top 20 candidates. The pool size matters: on real documents the right passage was in the top 8 only 45% of the time, and widening to 20 lifted full-hit from 0.45 to about 0.65.
Splits multi-part questions, searches in parallel, checks that every entity has evidence, and follows up. On real text it lifts multi-part full-hit from 0.36 to 0.72.
Models and what they cost
Cost per request is computed from provider-reported token counts and the price table in configuration. Models without a configured price show tokens only.
| Role | Default model | Setting |
|---|---|---|
| Answer | gpt-6-luna | CHAT_MODEL |
| Rerank | gpt-6-luna | RERANK_MODEL |
| Plan and reflect | gpt-6-luna | PLANNER_MODEL |
| Embeddings | text-embedding-3-small | EMBED_MODEL |
| Eval judge | same as answer model | JUDGE_MODEL |
Reasoning models spend hidden tokens and time. In the breakdown under an answer, reranking with gpt-6-luna is typically the largest single stage; a small non-reasoning model was faster and about as accurate on the real-document test (see Results).
Security: documents are untrusted input
Anything a user uploads ends up inside prompts, so a poisoned document can try to give the model orders. The threat model and the defences:
- Instructions hidden in a document (“ignore previous instructions…”).
- Breaking out of the source frame with fake closing tags or chat markup.
- Invisible characters used to slip past filters.
- Data exfiltration through an image or link the model is told to print.
- Manipulating the reranker or reflector, which also read passages.
- Passages are wrapped as data in every prompt, with explicit rules that sources are never instructions.
- Hidden characters are stripped and our own delimiter tags are neutralised.
- An ingest scan flags instruction-like documents (warning icon in the document list). It reports; it does not block.
- Answers cannot render images, and only plain web links are clickable.
- The model has no tools or actions, so the worst case is a wrong answer, not a side effect.
Honest limits: prompt-level defences are not guarantees, and the API has no user authentication, so do not expose an instance that holds private documents without an authentication layer in front of it. Measured attack results are on the Results page.