Transcribing Old Documents: Conventions, Uncertainty, and What to Leave Exactly as Written

Transcription conventions for old documents: what to preserve as written, how to expand abbreviations and mark uncertainty, and where machine drafts fit in an auditable workflow.

Leo Team

August 6, 2026

Transcribing Old Documents: Conventions, Uncertainty, and What to Leave Exactly as Written
Contents

Transcribing old documents is less about typing than about deciding, in advance, how literal your text will be. This is a working guide to the conventions that make a transcript auditable — what to preserve, what to expand, how to mark uncertainty, and where machine transcription belongs in the process.

Transcribing old documents well is not a matter of typing what a page says. It is a matter of deciding, in advance and in writing, how literal your text will be — what you preserve, what you expand, what you flag as uncertain, and how a reader can audit your text against the image. There is no single correct convention. What there is, across every serious tradition from the TEI Guidelines to the Library of Congress volunteer rules, is a shared requirement: declare your operations, keep the original reading recoverable, and never silently improve the source.

That last point separates a usable transcription from a decorative one. A transcript that quietly modernizes spelling, resolves an abbreviation without saying so, or repairs an apparent scribal slip has destroyed evidence that someone downstream — a paleographer, an editor, a title examiner, a genealogist checking a surname spelling — will need.

Decide what kind of transcript you are making

Before conventions, the deliverable. The same page can legitimately produce several different texts, and most of the disagreement people encounter in transcription guidance comes from projects using the same vocabulary for different products.

Diplomatic / documentary

Source-oriented. Preserves original spelling, abbreviation signs, punctuation, alterations, layout, lineation, and page or folio structure, so a reader can audit the text line by line against the image. The TEI Guidelines on primary sources distinguish a documentary approach that stresses the physical writing process from a textual approach that stresses textual alterations — and say plainly that which is appropriate depends on the project's needs. There is no default.

Semi-diplomatic

A working compromise, and in practice the most common scholarly choice. It keeps high-value source signals and makes a declared set of interventions. The Folger's Shakespeare Documented conventions are a concrete model: retain original spelling, i/j and u/v as written, ff, capitalization, punctuation, lineation, and indentation — but expand abbreviations, marking the supplied letters. Note the word "practical" in their own description. Semi-diplomatic is a reasoned editorial position, not a fixed standard you can cite as universal.

Normalized / reading text

Optimizes legibility, searching, indexing, or translation. Regularizes spelling and punctuation and usually flattens lineation. The Folger publishes modernized texts alongside the semi-diplomatic ones, which is the safer pattern: a reading layer is useful precisely because the literal layer and the image remain available beside it.

Pick one as your primary deliverable and say so at the top of the file. If you need two, produce two clearly labelled layers rather than a single hybrid that no one can interpret.

What to leave exactly as written

The general rule is narrow and firm: preserve anything that is graphic or orthographic evidence, and treat your instinct to correct as a signal to slow down.

  • Spelling as the writer spelled it, including inconsistent spellings of the same word or name on the same line. Early modern orthography was not standardized; a surname spelled three ways across three documents is data, not error.
  • i/j and u/v as written. "Iohn," "iustice," "vsed," "haue." The letter j was not part of the early modern English alphabet.
  • The long s (ſ), ligatures, thorn and yogh, doubled letters like ff, the ampersand or Tironian et, and superscript abbreviation marks. These are letterforms, not typos. They are also a documented weak point in machine transcription — a comparative study of historical OCR error patterns found systematic long-s→s and long-s→f confusions in both a fine-tuned model and a vision–language model, alongside implicit orthographic normalization and real-word substitutions.
  • Capitalization and punctuation as they appear, including punctuation that looks wrong. Both the US National Archives Citizen Archivist tips and the Library of Congress By the People rules direct volunteers to reproduce them.
  • Deletions, insertions, interlineations, marginalia, catchwords, and stamps. Struck-through text is often the most informative text on a legal page.

Where a modern reading matters — for searching, for a published index, for a client-facing summary — put it in a separate field or a linked alternative, not over the top of the original.

Abbreviation: expand, or record, or both

Abbreviation is where transcription becomes editorial. A brevigraph, a macron over a vowel standing for an omitted nasal, a superscript letter, a 9-sign — each is a mark that stands for absent letters, and each demands a decision.

Three defensible positions:

  1. Retain the mark. Most faithful, least readable, and hardest to search. Appropriate for paleographic and codicological work.
  2. Expand visibly, marking the supplied letters. The common scholarly compromise. English practice typically uses square brackets — `yo[u]r` — while the Newberry's Italian Paleography standards expand into round parentheses: `no(n)`. Both are legitimate; neither is universal, which is exactly why the legend matters.
  3. Record both. In TEI, keep the abbreviated form and the expansion as linked alternatives — `<choice>` wrapping `<abbr>` and `<expan>` — so the machine-readable record never erases what the scribe actually wrote, whatever the display layer shows.

Cross-language expansion policy does not transfer. English secretary hand, French notarial and chancery hands, German Kurrent and Sütterlin, Spanish procesal and cortesana, Dutch registers — these traditions have their own repertoires and their own project-specific policies, and no single authoritative cross-language rule set exists. Build a short expansion table for your corpus and language, and publish it with the transcript. Worth naming: German Kurrent is a German handwriting tradition. The script is the Latin alphabet; the language is German. Conflating a script family with the Latin language is a persistent and consequential confusion. If you are building that table from scratch, the survey of manuscript abbreviations and ligatures covers the main scribal shorthand types and where they recur across languages.

Marking uncertainty honestly

An unreadable word, an editorial conjecture, and a deliberate omission are three different things, and a transcript that renders all three as `[illegible]` has thrown away information a later reader needs.

TEI separates them:

  • `<unclear>` — a reading that cannot be transcribed with certainty.
  • `<gap>` — material omitted, with a recorded reason (illegible, damaged, editorially skipped) and extent.
  • `<supplied>` — text the editor supplies to fill damage or a scribal omission.

Record the reason, the approximate extent, your degree of certainty, and — on a multi-reader project — who made the call.

Visible sigla, meanwhile, genuinely diverge by community. NARA uses `[illegible]` and allows optional bracketed notations. The Library of Congress uses `[?]`. The Smithsonian Transcription Center uses `[[?]]` and double-bracket workflow labels such as `[[strikethrough]]` and `[[left margin]]`, and instructs volunteers not to type ditto marks but to write out what the ditto stands for. FromThePage leaves the convention to the project owner. These are workflow choices, not interchangeable semantics.

Nor should you borrow a bracket legend from another discipline without saying so. In the Leiden tradition inherited by epigraphy and papyrology, square brackets mean something specific: Bodard's account of EpiDoc notes that a pair of square brackets always signifies text lost through physical damage. Leiden+ documentation assigns angle brackets to text omitted by the scribe and supplied by a modern editor, and braces to surplus text the editor deletes. Import those brackets into an early modern will without a legend and you will mislead anyone who reads them literally.

The practical rule: publish a short legend at the head of the transcript. Half a dozen lines is enough.

Sic, and the temptation to correct

`<sic>` in TEI contains "text reproduced although apparently incorrect or inaccurate." That is all it does. It asserts that the odd reading is in the source; it does not mean you corrected anything. The correction, if you make one, is `<corr>`.

Two failure modes follow. The first is using sic as a badge on every unfamiliar form. The MLA Style Center is direct on the narrower case: if a spelling simply follows a different standard system, it is not an error and sic is not needed. Early modern spelling was not standardized at all, so tagging it would mean tagging most of the page.

The second failure mode is silent repair. Whether a scribal slip is retained bare, marked, or corrected in a reading text depends on your declared audience and the risk of misunderstanding — this is genuinely unsettled. What is not negotiable is that the original reading stays recoverable. In a documentary transcript, retaining the error plainly and noting it in a separate apparatus is usually cleaner than sprinkling sic through the text.

Layout: preserve it when it carries evidence

Whether to reproduce lineation is one of the clearest divergences in the published guidance, and it is not a disagreement about rigour. It is a disagreement about purpose.

The Library of Congress asks volunteers to preserve line breaks, on the grounds that they make review against the image easier. NARA explicitly says not to worry about matching the original format, including line and paragraph breaks, because their goal is searchable text. TEI provides `<lb/>` for a typographic line and `<pb/>` for a page boundary, and notes that these markers support proofreading and synchronization with page images.

Preserve line, page, and folio structure when physical arrangement carries meaning: paleographic analysis, scribal process, the placement of a legal clause, a metes-and-bounds call, marginal releases on a deed. Flatten it, if you like, in a labelled reading or index layer — but declare the flattening. And when you are working with columnar material, the column is the structure: a census page or a station register transcribed as running prose has lost its meaning, which is why reading census records is a column-by-column discipline rather than a linear one.

Where machine transcription fits — and where it must not

Almost everyone now begins with a machine draft. The conventions above do not become less important as a result; they become the standard against which the draft is judged.

The relevant question is not whether a model reads a page but what it does with the very features you have just committed to preserving. General-purpose chat models are the sharpest case. Trained overwhelmingly on modern text and working from a downsampled image, they tend to produce fluent, modern-looking prose. Archaic spelling gets regularized. The long s becomes s or f. An abbreviation mark is silently resolved, or dropped. The result reads smoothly and diverges from the page in exactly the places your conventions exist to protect. Garbled OCR announces its errors; a plausible substitution does not, which is why fluent LLM transcription errors are harder to catch than obviously broken output. Benchmark results in this area are also mixed and corpus-bound — the 2025 MLLM handwriting benchmark found accuracy varying by use case across a small set of datasets and languages, with the authors themselves flagging limited multilingual coverage — so no result there licenses skipping verification.

This is the stage Leo is built for. ATR-1, Leo's transcription model, reads Latin-script material — any language written in the Latin alphabet, handwritten or printed, roughly the past five hundred years — and it is trained to transcribe what is on the page rather than to normalize it. Strikethroughs, insertions, marginalia, tables, archaic orthography, and expansions in the form `yo[u]r` survive into the output rather than being smoothed away. On a randomized 97-image sample of early-modern English manuscripts from the Folger Shakespeare Library, ATR-1 scored roughly 5% character error rate at release, against approximately 13% for Transkribus's Text Titan I, 23.3% for Claude Opus, 24.8% for Gemini 2.5 Pro, and 56.7% for GPT-4.1 — the full comparison is published here. One corpus, one language, one moment in time; treat it as directional and test on your own material.

Two structural points matter more than the number. First, the base transcription is never overwritten. If you want a modernized reading, a translation, or a corrected version, those run as Transformations into separate tabs — so the literal layer and the interpretive layer stay distinct in the same way your file conventions require. Transcription and translation remain two jobs, not one. Second, there is no model to train first: nothing about the workflow asks you to key ground truth before you can read a page. Where the material is a pre-printed ledger or deed-book form with dense handwriting in the fields, expect weaker results — printed structure can dominate the manuscript entries — and plan review accordingly.

Machine output is a first pass. Errors of the recoverable kind — a wrong character, a wrong word — are caught by reading text beside image, which is why verifying transcription accuracy means prioritizing names, dates, sums, and boundary calls rather than sampling evenly.

A minimum legend to publish with any transcript

Six declarations, and you have made your text auditable:

  1. Deliverable type — diplomatic, semi-diplomatic, or reading text.
  2. Orthography — preserved as written, or normalized (and where the literal layer lives if so).
  3. Abbreviations — retained, expanded in square brackets, expanded in parentheses, or encoded as linked alternatives.
  4. Uncertainty tokens — your symbols for unreadable, conjectural, and omitted, defined separately.
  5. Layout — lineation and page/folio markers preserved or flattened.
  6. Sourcing — the image or facsimile each transcript corresponds to, and who transcribed and reviewed it.

None of this is exotic. It is the same discipline that governs the rest of a historical paleography practice: read the letterform before you read the word, and record what you saw rather than what you expected.

The unresolved question in this field is not which convention is correct. It is how to expose both a literal and a normalized text to readers and search engines without the normalized layer being mistaken for what the document says. The defensible answer is not a single style. It is provenance, linked alternatives, a published legend, and a text that anyone can hold up against the image and check.

Frequently Asked Questions

How do you transcribe old documents accurately?

Accurate transcription of old documents starts with a decision, not with typing: choose your deliverable — diplomatic, semi-diplomatic, or normalized reading text — and declare it at the top of the file. Preserve spelling as written, including i/j and u/v, the long s, ligatures, and abbreviation marks. Reproduce capitalization and punctuation even where they look wrong, and keep deletions, insertions, and marginalia. Mark unreadable words, editorial conjectures, and omissions as three separate things. Publish a short legend defining your symbols, and record which image each transcript corresponds to, so a reader can check your text against the page.

What is the difference between a diplomatic and a semi-diplomatic transcription?

A diplomatic transcription is source-oriented: it preserves original spelling, abbreviation signs, punctuation, alterations, layout, lineation, and page or folio structure so the text can be audited line by line against the image. A semi-diplomatic transcription keeps most of those source signals but makes a declared set of interventions — the Folger's Shakespeare Documented conventions, for example, retain original spelling, i/j and u/v, ff, capitalization, punctuation, and lineation while expanding abbreviations and marking the supplied letters. Semi-diplomatic is a reasoned editorial position rather than a universal standard, which is why you have to state which one you produced.

Should you correct spelling mistakes when transcribing historical documents?

No — not silently, and usually not at all in a source-oriented transcript. Early modern orthography was not standardized, so inconsistent spellings of the same word or surname are evidence, not error. Retain the reading as written and, if a correction matters, record it separately: TEI keeps `<sic>` for text reproduced although apparently incorrect and `<corr>` for the correction. Avoid using sic as a badge on every unfamiliar form; the MLA Style Center notes that a spelling following a different standard system is not an error. Whatever you decide, the original reading must stay recoverable.

What do square brackets mean in an old document transcription?

It depends on the tradition, which is why every transcript needs a legend. In much English scholarly practice, square brackets mark letters supplied when expanding an abbreviation — `yo[u]r`. The Newberry's Italian Paleography standards instead use round parentheses for expansions — `no(n)`. In the Leiden tradition used by epigraphy and papyrology, a pair of square brackets always signifies text lost through physical damage, with angle brackets for scribal omissions supplied by an editor and braces for surplus text deleted. Borrowing one discipline's brackets into another document type without explanation will mislead anyone reading them literally.

Can AI transcribe old handwriting, and how accurate is it?

Machine transcription is now a normal first pass, but its usefulness depends on whether the model preserves the features you have committed to preserving. General-purpose chat models, trained largely on modern text, tend to regularize archaic spelling, turn the long s into s or f, and silently resolve abbreviation marks — fluent output that quietly diverges from the page. Leo's ATR-1 model is trained to transcribe rather than normalize Latin-script material across roughly the past five hundred years. On a 97-image sample of early modern English manuscripts it scored about 5% character error rate at release, against roughly 13% for Transkribus's Text Titan I and far higher rates for general chat models. Treat any such figure as directional and verify against the image.

Share this article

© 2026 Leo Technologies Limited. All rights reserved