How answers are made
Every answer on this site is assembled the same way, under the same rules. This page describes those rules precisely enough that you can check us against them — because a citation you can't interrogate is decoration, and this whole product exists to be interrogated.
Your question is used to search the collection's records — nothing else goes into the answer. (The question itself is also kept for accuracy checking; “What is kept,” below, says exactly what that means.) The passages that come back are the only material the language model is given to write from. If the search returns nothing relevant, no model is called at all: you get a fixed “the records don't show this” response. An answer that does not rest on retrieved records cannot be served — that is enforced by the software, not promised by a policy.
Every answer must cite at least one record, and every citation is validated against the set of records actually retrieved for your question — the model cannot cite something it wasn't shown. If its draft fails that check, the draft is discarded and you get a plainer summary built directly from the retrieved passages instead.
Behind each numbered citation sits the record itself, and on the record page a panel shows the exact passage retrieved for your question — verbatim, never a summary. The sidebar is honest about the difference between retrieved and cited: every record the search returned is listed, but only records the answer actually rests on carry a number. A record listed without a number is one the model saw and declined to use — you can see what it saw.
By default, answers stay inside the records. On some collections the owner has enabled an experimental mode a reader can switch on per question, which permits the model to add historical context the records themselves don't contain. Three rules keep it honest:
- Two switches, both required. The collection's owner enables the capability; you invoke it per question. Neither alone does anything.
- Context is marked, sentence by sentence. Any sentence without a citation is rendered as a margin note labeled “The LLM's interpretation — uncited,” visually subordinate to the cited prose. The label is applied by the renderer, not by the model's good behavior — an uncited sentence cannot appear dressed as a claim about the records.
- The floor doesn't move. Even in this mode, an answer with no valid citations is not served. Interpretation is only ever a margin around the records, never a replacement for them.
A citation can resolve and still be undeserved — a sentence can cite a real record for something that record doesn't say. So on interpretive answers a second, independent model pass re-reads every cited sentence against the retrieved passages of the record it cites. A sentence the check cannot support is demoted: it loses its citation and is re-rendered as the LLM's own interpretation, labeled as such. Nothing is silently deleted, and the error direction is deliberate — authority is only ever understated, never overstated.
These are historical collections. They contain the vocabulary, the science, and the judgments of their moment — some of it wrong by today's scholarship, some of it painful to read. The product's job is to report what the records say and who says it, not to launder it into settled fact and not to sanitize it out of view. Cited text is what the record says; whether it was right is a judgment the product deliberately leaves with you.
Some collections have been given a register: their records separated out with their dates, the people in them and the places they name, the names and vocabulary reviewed by a scholar, the numbers pinned and checked against the collection at every deploy. On those shelves a question that counts, ranks, finds by date or name, or combines two of these is calculated from the register— the rule that produced the number is shown under the answer as How this was counted, and a reading of a computed set shows what it read under How this was read. No language model does the arithmetic. Where the register cannot support a question, the desk says so by name rather than estimating from a search; where a shelf has no register, counts are declined and the shelf says so before you ask. Which shelves have a register, and what each has been checked for, is stated on every shelf page under What this shelf can compute.
Every question asked here is kept for up to six months, to check answers for accuracy. A question is kept with: the collection it was asked on; the date (never the time of day); whether it was answered, refused, or answered from a finding aid; which model wrote the answer and at what reasoning setting; whether the question was routed through the collection's index; and how long the answer took to build. Nothing else is kept with it — not the answer, not any record text, and nothing about who asked: no account, no address, no identifier of any kind. Questions are deleted after six months — except any the developer has reviewed and kept as standing test questions, which stay indefinitely and still carry nothing about who asked.
Under every answer you can say whether it was helpful — “Yes” or “Not quite”, and if not, why: it was wrong, it couldn't answer, or it answered a different question. You are never asked whether an answer was correct: a fluent, well-cited answer can be wrong in ways a reader cannot see, so correctness is the developer's judgement, made later, not yours. A rating is noted with the question on the same terms as the question itself — the date, the collection, the rating and its reason, and which part of the system answered; no answer text, nothing about you; deleted after six months.
Two things keep an answer, and both are your choice. If you press “Something look wrong?” under an answer, that report — the answer, how it was built, and your note — is kept for up to two years so the problem can be checked; pressing send is the consent. If you tick keep this answer with my rating, the same kind of record is kept for the same time, whether you rated the answer helpful or not — a good answer kept is how the developer learns what works, not only what fails. Neither carries anything identifying you. Ratings and reports decide what the developer reads first; they never change how an answer is built.
Site administrators can record their own test conversations for debugging; this never applies to readers' sessions.
- A collection is what was digitized and ingested — no more. Absence of evidence here is absence from the collection, not from history.
- Retrieval is imperfect. The search can surface a thematically nearby record or miss the best passage inside a long document. The retrieved-versus-cited sidebar exists so you can see this happening.
- The checks are themselves model judgments. The faithfulness pass reduces miscitation; it does not prove correctness. The verbatim excerpt panel is there so the final check can always be yours.