Comparing Manuscript Versions: How to Collate Copies, Drafts, and Duplicate Registers From Transcribed Text

Comparing manuscript versions through a layered collation method: one faithful transcription per witness, normalization as a separate documented layer, then alignment so accidentals and variants remain recoverable.

Leo Team

September 17, 2026

Contents

Comparing manuscript versions is decided long before the collation runs — by what each transcription preserves and what it quietly regularizes. This article sets out the layered method: one faithful transcription per witness, normalization as a documented layer, then alignment, analysis, and interpretation. It also marks the points where the evidence you are looking for disappears without leaving a trace.

Comparing manuscript versions means collating witnesses — individual surviving copies, drafts, register entries, or prints that give evidence for a text — to find where they agree and where they diverge. The method that holds up is layered: keep one faithful, image-linked transcription per witness; add normalization as a separate, documented layer rather than editing it into the base; then align, analyze, and interpret. The order matters because a comparison layer can be rerun, while a destructively normalized base cannot recover the spelling, abbreviation, punctuation, or layout it discarded. Collation locates differences; it does not decide which reading is authorial, original, or legally authoritative.

What counts as a witness, and why the answer is wider than you think

For a literary scholar, the witness set is usually obvious: autograph draft, fair copy, scribal copies, first print. For historians the set is wider and stranger, because administrative systems produced parallel texts as a matter of routine.

An enrolled copy is a record entered into an official series — Chancery enrolments among them, described in The National Archives' catalogue. An exemplification is an officially authenticated transcript of a public record; a New York courts law-library explanation treats it as an authenticated copy of a public record. A duplicate register is a parallel copy maintained for institutional, ecclesiastical, or legal reasons — the bishop's transcript being the familiar English case, a contemporary copy of a parish register.

These are not merely redundant texts. A divergence between an original and its enrolment or transcript can document copying practice, deliberate amendment, omission, authentication procedure, or simply the different evidentiary function the two objects served. What it does not do is settle authority. The legal effect of a certified copy is jurisdiction- and context-specific, and scholarship on notarial registration and certified copies — including comparative work on notarial systems — treats the question as historically variable rather than settled. If your argument depends on one version being authoritative, that has to be argued from the record series and the law of its time, not from the collation output.

Decide your transcription level before you transcribe anything

Collation is only as good as the texts fed into it, and the single decision that governs everything downstream is transcription level.

A strict diplomatic transcription reproduces what the source displays: spelling, punctuation, capitalization, word division, variant letter forms, layout, unexpanded abbreviations, and — at the strictest end — uncorrected slips of the pen. A semi-diplomatic transcription regularizes selected features under a declared policy. A normalized or modernized transcription keeps the words while updating surface forms. A critical text is an editorial construction that selects, orders, and emends readings across witnesses.

Driscoll's account of these levels, in the TEI's Electronic Textual Editing, is the standard reference and makes the crucial separation explicit: recording what is in the source is a different act from proposing what ought to have been there. All levels are legitimate for some purpose. The failure is not choosing a level; it is choosing one implicitly and then discovering that your base text no longer records what the document shows.

The loss begins earlier than most researchers expect. Regularizing `publick` to `public`, expanding a suspension mark, modernizing punctuation, changing case, or collapsing historical `u/v` and `i/j` distinctions can erase evidence even when the lexical word looks unchanged. Two illustrative failure modes:

  • False negative. Witness A reads `publick`; Witness B reads `public`. A faithful comparison reports an orthographic variant. If both bases were regularized to `public`, the variant simply is not there to find.
  • False positive. Both images read `ye`. One transcription pass expands it to `the`; another leaves it. The alignment now records a difference in editorial behaviour and presents it as a difference between documents.

Neither is exotic. Both are what happens when normalization is applied unevenly and invisibly.

Substantives, accidentals, and the temptation to discard the second

W. W. Greg's distinction between substantives (the words and their meaning) and accidentals (spelling, punctuation, capitalization, word division, presentation) is well established and genuinely useful. It is not a licence to discard accidentals.

Accidentals are frequently the evidence historians most need. Spelling and abbreviation practice can identify a scribe or a copying milieu. Punctuation and word division can date a hand or localize it regionally. A pattern of shared accidental readings across witnesses is standard evidence for filiation — for which copy descends from which. If you have flattened the accidentals to make the comparison "cleaner," you have removed the layer that would have told you how these documents are related to one another.

The conventional scribal-error vocabulary sits on the same footing. Haplography (writing once what should occur twice), dittography (unintended repetition), and homoeoteleuton (eye-skip between similar endings) each name a hypothesis about a copying process. They are diagnoses to be tested against the image and the other witnesses, not labels to be applied because the phrase fits.

Genetic material: drafts are not defective copies

Where your witnesses are authorial drafts rather than administrative copies, the question changes. Genetic criticism studies the writing process and its successive states — cancellations, additions, rearrangements, revisions — rather than reconstructing a single preferred text. The French term avant-texte covers the documents and states leading toward a work.

Ferrer's essay in Variants is worth reading in full before you commit to an apparatus, because it sets out both the difference in aims and the genuine limit: textual criticism tends toward establishing a text by eliminating variants, genetic criticism toward destabilizing it by confronting it with its states — and distinguishing creative variants from variants of transmission is sometimes impossible. That is not a reason to avoid the distinction. It is a reason to encode your evidence so that a later reader can disagree with your classification without re-transcribing the manuscripts.

Working editions show how this looks in practice across languages. Jane Austen's Fiction Manuscripts declares its transcriptions diplomatic and faithful to spelling, paragraphing, and punctuation. The German Faustedition combines manuscript archive, text-critical print, and visualizations of genesis. The critical-genetic edition of Ariosto's autograph fragments works from 58 surviving leaves — 54 in Ferrara, two in the Ambrosiana, two in Naples — and documents its choices explicitly: original word joining, capitalization where discernible, punctuation, and original i/j distinctions preserved, with selected conventions declared. That project is a useful corrective to a common assumption: "diplomatic" still involves declared, purpose-specific policy. There is no single universal level of diplomatic detail.

The layered architecture that survives review

The defensible pipeline for comparing manuscript versions runs in this order:

  1. Image or facsimile, retained and citable.
  2. One faithful, image-linked transcription per witness, at a declared level.
  3. Documentary or genetic encoding — additions, deletions, substitutions.
  4. An optional normalization layer, explicitly documented and reversible.
  5. Tokenization.
  6. Alignment.
  7. Variant analysis.
  8. Visualization.
  9. Human editorial interpretation.

Steps 4 through 8 are rerunnable. Step 2 is not, if you got it wrong.

CollateX documents a five-part Gothenburg-style workflow — tokenization, optional normalization, alignment, analysis, visualization — and is designed around that separation of concerns. Its output can be represented as a variant graph in which shared segments merge and differences branch; the technical discussion at Balisage sets out the representation in detail. Two of its own caveats belong in your method note rather than buried in documentation: the built-in normalization options are deliberately simple, and heuristic transposition detection cannot guarantee correctness given external evidence. Treat CollateX as a candidate-finder and a visualizer, not an adjudicator.

Juxta Commons, TUSTEP/TXSTEP, Stemmaweb, and TRAViz cover adjacent functions — pairwise comparison, textual-data processing, stemmatic analysis, variant-graph visualization. Choose by the question, not the interface.

Encoding: what `<lem>` does and does not mean

For interoperable output, the TEI's critical-apparatus module gives you `<app>` for an apparatus entry, `<lem>` for the lemma or base reading, and `<rdg>` for an alternative reading. It also distinguishes location-referenced, double-end-point, and parallel-segmentation methods — which encode genuinely different relationships between base and variants and should not be conflated. Double-end-point attachment marks both ends of the lemma, permitting unambiguous matching; parallel segmentation expresses all variants at a point as variants on one another, so no two variations may overlap.

For the documentary side, `<sourceDoc>` keeps the transcription tied to the physical object, and `<add>`, `<del>`, and `<subst>` record interventions on the page — the machinery described in the chapter on representing primary sources.

One misreading to avoid: `<lem>` is not a claim that a reading is authorial or historically correct. It records the lemma selected for that apparatus. Selecting it is an editorial act, and TEI standardizes representation rather than settling a theory of textual authority. Copy-text editing, eclectic reconstruction, documentary editing, genetic editing, and New Philology's emphasis on the plurality of witnesses all answer different questions; the MLA Guidelines for Editors of Scholarly Editions is the place to argue out which you are following. Say which, in the edition.

Getting the base transcriptions made

The bottleneck is obvious once the architecture is clear. Collation software is cheap and fast; producing a faithful transcription of every witness is neither. A multi-witness project is a multiplication problem — the Middle Dutch diplomatic dataset covering Dietsche Catoen (16 witnesses), samples of the Rhymed Bible/Scolastica (15 witnesses), and Karel ende Elegast (14 witnesses) amounts to almost 29,000 verses of diplomatic transcription. That is one project's figure, not a general rate, but it indicates the scale of the labour behind any serious collation.

Machine transcription can take the first pass, and for a multi-witness comparison it changes what is feasible. It also introduces the specific risk at issue here: a transcription engine that silently normalizes destroys variants before you ever see them. General-purpose chat models are the sharpest version of the problem, because fluent output is plausible output — an expanded abbreviation or a modernized spelling reads as correct and is very hard to catch downstream. There is no controlled cross-language benchmark establishing that every such model normalizes every historical text, and it would be overclaiming to say so; the point is that the failure mode is invisible at exactly the level collation depends on, so verification against the image is not optional.

This is the stage Leo is built for, and it is why source integrity is the design commitment rather than a feature line. Leo's ATR-1 model transcribes what is on the page: strikethroughs, additions, margin notes, archaic orthography, and abbreviations left as written rather than silently resolved. It reads Latin-script material — any language written in the Latin alphabet, so English wills, French notarial registers, Dutch registers, German parish books, Italian and Spanish witnesses all sit in scope — and it runs zero-shot, with no per-corpus model training before you can read your first witness. Non-Latin scripts (Greek, Cyrillic, Hebrew, Arabic, Indic, East Asian) are out of scope.

Three things matter specifically for version comparison. First, retranscription writes to a new tab rather than overwriting, so the base transcription of each witness stays intact while you work above it. Second, the AI Transformations — Correct, Modernize, Translate, Interpolate — each write their output to a separate tab: that is normalization as a documented layer, structurally rather than by discipline. If you want a normalized comparison text, you generate it from the base and both survive. Third, export includes TEI XML, so the faithful base can carry into the apparatus encoding you have chosen rather than being retyped into it.

Leo does not collate. Alignment, variant graphs, and stemmatic analysis belong to CollateX and its neighbours, and Leo's position is upstream of them: it produces and manages the per-witness transcriptions, the metadata, and the image-beside-text record that collation consumes. Errors it does make are the recoverable kind — a wrong character or word, checked against the image displayed beside the text — and its output checks hide and retry suspect results rather than presenting them as finished. That is a mechanism, not a guarantee; every reading still needs a human eye where the argument rests on it.

Provenance is the deliverable

Whatever your tooling, the thing a reviewer or a later editor needs is the audit trail. Retain the scan; the raw machine output; the correction decisions; the declared transcription level; the normalization rules; the software and version settings; and the derived comparison. Character error rate, defined as edit distance from a ground-truth transcription normalized by reference length, is a recognition metric — it tells you nothing about whether a system quietly modernized a spelling or expanded a suspension mark in a way that dissolved a variant. Ground truth is itself a policy artefact, which is why projects publishing it, such as OCR-D with its ground-truth transcription guidelines, document their transcription rules alongside the text.

None of this is only a Latin-language problem, and it is worth saying because the phrase "manuscript variants" still summons classical stemmatics. The relevant category is Latin script. English, French, German, Dutch, Italian, and Spanish witnesses each have their own histories of spelling variation, abbreviation, punctuation, and word division; the Spanish standards literature on diplomatic transcription identifies respect for original spelling as a general principle, and each tradition needs its own rules for abbreviation marks, letter forms, and capitalization. The method that travels is not identical character treatment. It is transparent preservation of the documentary state before interpretation — the same principle that governs a whole historical research workflow, from capture through citation.

The discipline this asks of you is modest and mostly procedural: decide the level, declare it, keep the image beside the text, and put every act of regularization somewhere a reader can see and undo it. Do that, and the interesting work stays available — the divergence between an original and its enrolment that reveals a clerk's amendment, the shared misspelling that groups three copies against a fourth, the cancelled phrase that shows a draft changing its mind. Flatten the base, and those findings do not become wrong. They become invisible, which is worse, because nothing in the output will tell you they were ever there.

Frequently Asked Questions

How do you compare manuscript versions of the same text?

Comparing manuscript versions works best as a layered process. Keep one faithful, image-linked transcription per witness at a declared level; add any normalization as a separate, documented layer rather than editing it into the base; then tokenize, align, analyze the variants, visualize, and interpret. The order matters because the comparison layers can be rerun, while a destructively normalized base cannot recover the spelling, abbreviation, punctuation, or layout it discarded. Collation locates where witnesses agree and diverge. It does not decide which reading is authorial, original, or legally authoritative — that argument comes from you.

What is the difference between a diplomatic and a normalized transcription?

A strict diplomatic transcription reproduces what the source displays: spelling, punctuation, capitalization, word division, variant letter forms, layout, unexpanded abbreviations, and at the strictest end uncorrected slips of the pen. A semi-diplomatic transcription regularizes selected features under a declared policy. A normalized or modernized transcription keeps the words but updates surface forms. A critical text is an editorial construction that selects and emends readings across witnesses. All are legitimate for some purpose. The failure is choosing a level implicitly, then finding the base text no longer records what the document actually shows.

Can AI transcription tools be used for collating manuscript witnesses?

Machine transcription can take the first pass and makes multi-witness comparison feasible, but the engine must not silently normalize. An expanded abbreviation or modernized spelling reads as plausible and is very hard to catch downstream, and it removes the variant before you ever see it. Leo's ATR-1 model transcribes what is on the page — strikethroughs, additions, margin notes, archaic orthography, abbreviations left unresolved — across Latin-script languages, and exports TEI XML. Leo does not collate; alignment and stemmatic analysis belong to tools like CollateX. Every reading your argument rests on still needs checking against the image.

What counts as a manuscript witness — do enrolled copies and duplicate registers qualify?

Yes. A witness is any surviving copy, draft, register entry, or print that gives evidence for a text, and administrative systems produced parallel texts routinely. An enrolled copy is a record entered into an official series. An exemplification is an officially authenticated transcript of a public record. A duplicate register is a parallel copy maintained for institutional, ecclesiastical, or legal reasons — the bishop's transcript of a parish register being the familiar English case. Divergence between an original and its enrolment can document copying practice, deliberate amendment, omission, or authentication procedure. It does not by itself settle which version is authoritative.

Why do spelling and punctuation differences matter in collation?

Because accidentals — spelling, punctuation, capitalization, word division, presentation — are frequently the evidence historians most need. Spelling and abbreviation practice can identify a scribe or a copying milieu. Punctuation and word division can date or localize a hand. A pattern of shared accidental readings across witnesses is standard evidence for filiation, showing which copy descends from which. Flattening accidentals to make a comparison look cleaner removes the layer that would have shown how the documents relate. Greg's distinction between substantives and accidentals is useful for analysis; it is not a licence to discard the second.

Share this article

© 2026 Leo Technologies Limited. All rights reserved