About

Ask the Records is a reading desk for historical collections. You pick a collection, ask questions in plain language, and get answers drawn from that collection's own digitized records — every answer cites the records it used, and every record points back to its source. It is a way to interrogate an archive, not a chatbot that claims to know history.

How it works

Two mechanisms, and each answer says which one made it. Where a collection has a register — its records separated out with their dates, names and places, reviewed — a question that counts, finds by date or name, or combines is calculated from that register, with the rule shown under the answer; no language model does the arithmetic. Every other question is answered from the passages themselves: the system searches the collection's records for the passages most relevant to it, composes an answer from those passages alone, and cites each record it drew on. If nothing relevant is retrieved, it says so instead of improvising. Nothing is trained on; nothing is inferred beyond what the records hold.

The full rules — how citations are validated and re-checked, what the labeled interpretation mode is, and where the limits are — are written out at How answers are made.

The commitments

These are the distinctions the archives community now writes policy about, and they are commitments here, not features:

  • The records are never trained on. They are read fresh for every question; no model learns them, and your material is consulted, not absorbed.
  • Deleted means deleted. Removing a collection removes its records and its search index.
  • Every answer cites the records it uses — and behind each citation you can read the exact passage the answer drew on, verbatim. A citation here is checkable, not decorative.
  • We don't impersonate historical figures. Asking “the Frederick Douglass Papers” means asking the archive that holds his words; the answer speaks about the records in the third person, never in his voice.
  • We don't claim completeness. A collection is what has been digitized and ingested, no more. When the records don't show something, the answer says so rather than filling the gap.
Who made it

Ask the Records was built by Kevin Hegg, a digital humanities practitioner with a master's in History who works in an academic library at the intersection of digital scholarship and special collections. It is built the way that vantage point demands: provenance first, claims checkable, and the technology kept in service of the records rather than the other way around.

For institutions

Ask the Records is a working prototype, not a product. It runs on a single machine and holds six collections. I built it to find out whether the approach is useful to people who work with or research in the archives.

It is not a general-purpose chatbot. Every answer cites the records used to construct it. Where a collection has been given a register, that register is built from the archive's own structure — its correspondence, its attendance lists, its agenda — and the counts and lookups come from it deterministically. Where the system cannot settle something, it says so instead of guessing.

At this point, I am not selling anything or adding new collections. I am asking for your critical assessment: Do the answers meet your scholarly and research expectations? Is it useful as a tool for interrogating and understanding the archival artifacts? Where does it need to be improved? Would you put your collections in the system if you were confident your content would not be used to train an LLM?

Contact

Please contact me at kevinhegg@gmail.com with comments and questions.

← Back to the collections