Citing Primary Sources: A Working Method for Archival References You Can Defend

Working method for defensible archival citations: the retrieval coordinates a footnote must contain, British versus American conventions, digital surrogates, and why transcription is itself an editorial claim.

Leo Team

August 18, 2026

Citing Primary Sources: A Working Method for Archival References You Can Defend
Contents

This is a practical guide to citing primary sources from archives — what a footnote has to contain, how British and American conventions differ, how to handle digital surrogates, and why the transcription behind a quotation is itself an editorial claim. It matters because a citation that cannot be followed back to the page is not a citation; it is a gesture, and it is the kind of thing that surfaces at proof stage when the volume is a thousand miles away.

A defensible archival citation is a set of coordinates, not a title and an institution. It should let another reader — or you, six years later, checking a proof — retrieve the exact page you read. At minimum that means the repository, the collection or fonds, the series or record group, the physical unit (box, volume, file, piece), the item-level identifier, and an internal locator such as a folio, page, entry or act number. Where you consulted a digital surrogate rather than the original, the citation should say so and record a persistent identifier if the repository supplies one. Everything else — which style manual, which punctuation — is downstream of having captured those coordinates in the first place.

That last point is the one that actually breaks people's footnotes. Style can be fixed at the copy-editing stage. A folder of unlabelled photographs from a trip you cannot afford to repeat cannot.

Why the coordinates, and not the title, do the work

Archival description is hierarchical. ISAD(G), the international standard, describes records at multiple levels, conventionally fonds or collection, then series, sub-series, file and item. DACS, the North American content standard, recognizes that four-level framework while allowing for other schemes of arrangement. A citation is, in effect, a path through that hierarchy, narrowed until only one page remains.

Each element earns its place:

  • Repository — the custodian, and the authority under which the record can be consulted.
  • Collection or fonds — the aggregate, and with it the provenance.
  • Series or record group — the record-creating activity or the arrangement imposed on it.
  • Box, volume, file, piece — the unit you physically order.
  • Item — the letter, act, docket, register entry, or manuscript number.
  • Folio, page, recto/verso, entry number — the passage itself.

Drop any one of these and the citation degrades from a retrieval instruction into a gesture. "Somerset Record Office, wills" is a gesture. So, in a different way, is a box number alone: the U.S. National Archives cautions against citing record group and box numbers on their own, and its guidance on citing federal records asks for record item, file unit, series, subgroup, record group and repository.

Two national traditions, and what they have in common

British and American practice differ in emphasis, and it is worth knowing which one your material speaks.

The UK tradition is reference-code-led. The National Archives asks for three things in a brief citation: the name of the institution, the full catalogue reference, and an internal identifier — "the folio, page, docket, membrane or other number within a box, volume, file, bundle or roll." The catalogue reference must be set out exactly as it appears in Discovery, with the same spacing and punctuation: `ADM 22`, `CO 700/MaltaandIonianIslands10a`, `LAB 2/10/L169/1901`. You normally cite the lowest available level in the catalogue hierarchy, then a comma, then the internal locator — `JUST 1/40, pp. 143–149`. The code is doing the hierarchical work; the comma marks where the catalogue stops and the volume's own foliation begins.

The American tradition is descriptive and hierarchical: the coordinates are spelled out rather than compressed into an alphanumeric string. Electronic records normally add an electronic-record designation and the online location.

Both traditions agree on the principle that matters: reproduce the repository's own identifier exactly, and never substitute your own description for it. A cote is not translatable. A Signatur is not paraphrasable. If the register you used is `1 E 234` in a French departmental archive, that string is the retrieval key, and rendering it as "notarial register, 1 E series" makes it unfindable.

The style manuals, briefly

Style guides govern presentation, not substance, and they disagree less than researchers fear.

MLA 9 treats the collection as a container: author or creator, then title or description of the work, then the collection as container title, then a location element carrying the institution and any box, file and manuscript numbers. It also allows descriptive in-text treatment, with no works-cited entry, for material in an unprocessed collection.

MHRA's Style Guide, Fourth Edition (2024), §7.7, gives the UK-style sequence: place, repository, collection or shelfmark, manuscript identifier, folio reference — full repository and collection names on first use, abbreviations thereafter, with fol./fols and recto/verso notation.

Chicago locates its manuscript-collection examples in chapter 14, where full identification of unpublished material usually runs from the item through the series title, the collection, and the depository. Chicago also holds that unpublished manuscripts unavailable for others to consult may be mentioned in the text or a note but are not usually given a bibliography or reference-list entry — a rule that surprises people the first time they meet it, and one worth checking in the edition your press actually uses. If you need a punctuation-by-punctuation comparison across Chicago 17 and 18, MLA 9 and MHRA 4, read the manuals themselves rather than reconstructing the differences from quick guides; the emphases are well documented, the edition-level changes less so.

None of these creates a universal international order, and none can. Which is why the safe habit is to record more coordinates than any single style requires, and let the style guide decide what to drop.

Vernacular records: preserve the local identifier

Much of the research historians do is in vernacular series — English wills, French notarial minutes, Dutch registers, German parish books — and each national system identifies the item its own way. There is no single European formula. The stable principle is that whatever the local identifier is, it goes in the citation unaltered:

  • English will or probate inventory — court or registry, series or jurisdiction, piece or file, testator, and probate date or folio/page.
  • French notarial act — archive, fonds or series, cote, notary or étude, date, act or register, folio.
  • Italian notarial act — state archive and fondo, notary, protocollo or register, date, act or folio.
  • German parish or civil register — archive Bestand and Signatur, parish or civil office, register type, year, page or entry.
  • Dutch register — archive and fonds, inventarisnummer or stuk, register and date, folio or entry.
  • Spanish protocolo — archive, fondo, signatura, notary and protocol, year, act or folio.

Treat these as practical patterns to verify against the holding repository, not as laws. Where a repository publishes its own explanation of its numbering, use it — French departmental archives publish notarial répertoires explaining cote structure, Italy's Antenati portal explains the organization of notarial archives in the state archives, and Spain's PARES portal documents its own conventions. Verify locally; retain original spelling and diacritics.

The digital layer: surrogates, PIDs, and rot

Increasingly you are not citing the manuscript. You are citing an image of it, and the two are not the same object.

A digital surrogate is a representation. It should be described as the representation you consulted, not silently conflated with the original. Where both provenance and the consulted representation matter — which is most of the time in serious work — cite both: the physical coordinates identify the object, the digital reference identifies what you actually read.

Prefer a persistent identifier over a viewer URL. Gallica, for instance, builds its referencing on ARK, and exposes it at two grains: the document (`https://gallica.bnf.fr/ark:/12148/bpt6k3228953`) and the digitized page (`.../f34`). Page-grain identity is what a footnote needs. Where a repository publishes IIIF, the Presentation API gives you a manifest for the work and canvases for individual pages or views, and the Content State API can identify a region of a canvas — finer targeting than a flat item URL. IIIF supplies the technical model; it does not yet supply a settled humanities citation convention, so record the identifier and let your discipline's conventions format it.

The reason to care is measurable, at least on the web side. Klein et al., in Scholarly Context Not Found (PLOS ONE, 2014), found that one in five STM articles suffered reference rot — the cited web context could no longer be revisited — and distinguished link rot (the URI is dead) from content drift (the URI resolves, but no longer returns what was cited). Zittrain, Albert and Lessig's Harvard Law Review study reported reference rot in more than 70% of URLs in the legal-journal material studied and in 50% of URLs in U.S. Supreme Court opinions. These are web-citation populations, not archival ones; there is no verified equivalent figure for archival citations, and none should be invented. But the mechanism is the same in both, and a PID mitigates it without eliminating it: an identifier improves resolution and identity, it does not guarantee that the image, viewer or transcription behind it stays fixed. Where the version matters, archive the representation you consulted.

The transcription is not the source

This is the failure that does the quiet damage in published work.

A transcription is an intervention. Diplomatic transcription preserves surface features — spelling, punctuation, capitalization, word division, variant letterforms, line division, unexpanded abbreviations, and in the strictest form even apparent slips of the pen, as Driscoll's account of levels of transcription sets out. Normalized transcription regularizes some of that. Both are legitimate; the choice is an editorial claim, and it must be disclosed in the edition or apparatus rather than hidden inside an apparently literal quotation. If your footnote presents modernized spelling as if it were what the scribe wrote, you have made a silent argument on the reader's behalf.

Two practical consequences follow.

State the level

Say whether a quoted passage is diplomatic, semi-diplomatic or normalized, and say it once, clearly, in your editorial note — then be consistent. Where you expand an abbreviation, mark the expansion (`yo[u]r`), and keep the mark in the citation-bearing text rather than resolving it away. Our guide to transcription conventions and marking uncertainty works through the notation in detail, and the reasoning behind preserving manuscript abbreviations and ligatures rather than silently expanding them is the same reasoning.

Keep the transcription anchored to the image

TEI provides the mechanism: `@facs` associates any element in a transcription with an image of the corresponding part of the source, `<facsimile>` and `<sourceDoc>` hold the digital facsimile, and `<msDesc>` describes the identifiable manuscript or other text-bearing object. Even if you never publish TEI, the discipline it enforces — every reading attached to a specific surface — is the discipline that makes a citation checkable.

Capture-time practice: the only stage you cannot redo

Metadata cannot be reliably reconstructed after the research trip. A folder of unlabelled images cannot be mapped back to boxes, files and folios with any confidence, and the point at which you notice is usually the point at which you are writing the footnote.

Columbia's Rare Book & Manuscript Library recommends the simplest fix: photograph the folder cover, which carries the reference, before photographing its contents, so that every time a folder cover appears in the run of images you know a new reference has begun. Sound institutional practice, not a quantified standard — and worth adopting anyway.

Build the rest around it:

  1. Shoot the reference before the contents — folder cover, spine label, or a slip with the code written on it.
  2. Shoot the opening in a way that preserves foliation: capture stamped or written folio numbers, and note recto/verso where the volume does not make it obvious.
  3. Record, at consultation time, the reference code, folder or volume label, folio/page range, date, document type, and any digital identifier.
  4. Note the access date and, for online material, the PID.
  5. Keep a running sequence so images stay in archival order.

A related discipline governs image quality itself: legible geometry, even lighting, and consistent focus are what make a page re-readable at all, which our field method for photographing archival documents covers. This capture stage is the first of the five that make up a working historical research workflow — capture, organize, transcribe, analyze, cite — and each one constrains the next. A reference lost at capture cannot be recovered at citation.

Where the tools sit, and where they don't quite meet

No single system currently guarantees a durable relation among repository code, image, folio, transcription, editorial status and eventual footnote. It is worth being clear about who does what.

Reference managers handle the citation itself. Zotero can be bent to archival material: Harvard's guide to Zotero for archival research recommends the Manuscript item type for a collection as a whole, the closest-matching type for individual items, the Archive field for the full repository name, and — importantly — the "Loc. in Archive" field for call number, box and folder plus the collection title. EndNote has a Manuscript type with Author, Title, Collection Title and Manuscript Number, but no native fonds-to-folio-to-image graph. Neither models archival hierarchy properly; both can be made to hold it if you are disciplined about which field carries what. Our note on where transcribed text fits in a Zotero-first workflow works through the field mapping.

Image managers handle the photographs. Tropy is built for organizing and describing research-material photographs, with custom templates and exports to JSON-LD, CSV and Omeka S — but as its developers state plainly on the Tropy forums, it is not a citation manager and does not generate citations. Omeka S supports flexible resource templates and linked internal resources, though the provenance schema is yours to design. DEVONthink is a general document manager, not an archival-citation standard.

None of them reads the page. That is the gap: you can have well-described images and still be unable to search, quote or verify a word of what is on them.

Reading the page at scale, without losing the coordinates

The stage where citation discipline most often collapses is transcription, because it takes the longest and is the most tempting to shortcut. A researcher with four hundred photographs of a notarial series can either spend a term reading them or hand them to something that reads faster — and the shortcut most people reach for first, a general chatbot, is the one that damages citations most, because it produces fluent, plausible readings that silently modernize spelling and resolve abbreviations. A misread that looks wrong is recoverable at proof stage. A misread that reads like clean early modern English is not, and it will be quoted verbatim in your article with a footnote pointing at a folio that does not say it.

This is the stage Leo is built for. It transcribes what is on the page rather than smoothing it: strikethroughs, interlinear additions, margin notes, tables, editorial expansions and archaic orthography survive the transcription rather than being tidied into modern prose — which is precisely what a diplomatic or semi-diplomatic citation needs. ATR-1, the model behind it, reads Latin-script material — any language written in the Latin alphabet, so English wills, French notarial minutes, Dutch registers, German parish books and Spanish protocolos are all in scope, along with printed matter of any period. Non-Latin scripts (Greek, Cyrillic, Hebrew, Arabic, Indic, East Asian) are not. Transcription is not translation: rendering a source into English is a separate Transformation that writes to a new tab, leaving the base transcription in the source language intact — which is the version your footnote should quote.

For citation purposes, two things matter more than speed. First, the image stays beside the text, so every reading can be checked against the surface it came from. Second, each document carries a structured metadata record — Title, Creator, Date, Archive, Type, Collection, Box, Folder, Identifier, Rights — so the coordinates you captured in the reading room stay attached to the pages rather than living in a separate spreadsheet. Export runs to TEI XML alongside Word, PDF and HTML, which is the route into a `@facs`-linked edition if that is where the project is going. What Leo does not do is generate formatted citations or manage your bibliography; that remains the reference manager's job.

The habit that makes it defensible

Citation is not a formatting task performed at the end. It is a chain of custody maintained from the moment you order the box: the repository's own code recorded exactly, the folio noted while the volume is open, the image kept beside the reading, the editorial level declared rather than assumed, and the digital identifier preserved where one exists. Every link in that chain is cheap to maintain at the time and expensive to reconstruct later.

The test to apply to any footnote you write is simple and slightly uncomfortable: hand it to a colleague who has never seen the collection, and ask whether they could order the volume and find the line. If they could, the citation is doing its job. If they could not, no style manual will save it.

Frequently Asked Questions

How do you cite primary sources from an archive in a footnote?

Cite a set of coordinates, not a title. A defensible archival footnote records the repository, the collection or fonds, the series or record group, the physical unit (box, volume, file, piece), the item-level identifier, and an internal locator such as a folio, page, entry or act number. If you consulted a digital surrogate rather than the original, say so and record a persistent identifier where the repository supplies one. Style — Chicago, MLA, MHRA — governs punctuation and order, and can be fixed at copy-editing. The coordinates cannot be reconstructed later, so capture more of them than any single manual requires.

What is the difference between British and American archival citation?

The UK tradition is reference-code-led; the American tradition is descriptive and hierarchical. The National Archives asks for the institution's name, the full catalogue reference reproduced exactly as it appears in Discovery — same spacing and punctuation — and an internal identifier such as a folio, page, docket or membrane number, usually cited at the lowest available catalogue level. American practice spells the hierarchy out instead of compressing it into an alphanumeric string, and the U.S. National Archives asks for record item, file unit, series, subgroup, record group and repository. Both agree on the essential point: reproduce the repository's own identifier without paraphrasing it.

How do you cite a digital image of a manuscript?

Cite both the physical object and the representation you actually read. A digital surrogate is a representation, not the original, and conflating the two hides what was consulted. Give the physical coordinates — repository, collection, series, unit, item, folio — then the digital reference, preferring a persistent identifier over a viewer URL. Gallica's ARK identifiers, for example, work at both document and page grain, and page-grain identity is what a footnote needs. Record the access date. A persistent identifier improves resolution but does not guarantee the image or transcription behind it stays fixed, so archive the version where it matters.

Why does it matter whether a transcription is diplomatic or normalized?

Because the transcription is an editorial claim, not a neutral copy of the source. Diplomatic transcription preserves spelling, punctuation, capitalization, word division, line division and unexpanded abbreviations; normalized transcription regularizes some of that. Both are legitimate, but presenting modernized spelling as if it were what the scribe wrote makes a silent argument on the reader's behalf. State the level once, clearly, in an editorial note and stay consistent. Mark expansions — `yo[u]r` — and keep the marking in the quoted text rather than resolving it away, so the footnote and the page agree.

Can Zotero or Tropy produce archival citations?

Zotero can hold archival references if you are disciplined about field mapping; Tropy cannot generate citations at all, as its developers state plainly. Harvard's guidance suggests the Manuscript item type for a collection, the Archive field for the full repository name, and "Loc. in Archive" for call number, box, folder and collection title. EndNote has a Manuscript type but no fonds-to-folio-to-image structure. Tropy organizes and describes research photographs and exports to JSON-LD, CSV and Omeka S, but it is an image manager. Neither models archival hierarchy natively, and neither reads what is on the page.

Share this article

© 2026 Leo Technologies Limited. All rights reserved