How to Test AI Transcription on Your Own Documents: An Afternoon Protocol
How researchers can test AI transcription tools on their own documents in an afternoon via random sampling, reference keying, CER scoring, and manual error reading for a directional local benchmark.
Leo Team
August 4, 2026

Contents
Published accuracy figures describe someone else's corpus, not yours. This is a step-by-step method for how to test AI transcription on your own documents — sampling, keying a reference, scoring character error rate, and reading the errors by hand — in a single afternoon. The result is a directional local benchmark you can defend in a footnote.
You can run a defensible tool comparison on your own manuscripts in one sitting. The method is short: pick a small random sample from the collection you actually intend to transcribe, key a careful reference transcription for those pages under one written convention, run every candidate tool on the identical images, score character error rate (CER) and word error rate (WER) with one scorer under one normalization policy, and then read the errors by hand to see what kind of mistakes each tool makes. The number tells you which tool is closer. The error categories tell you which tool is safe. Both matter, and the second matters more.
What you cannot do in an afternoon is produce a collection-wide accuracy figure with narrow uncertainty. Ten or twenty pages give you a directional local benchmark for your material — which is exactly the question a researcher choosing a tool needs answered.
Why vendor and published figures don't answer your question
Every published error rate is a statement about a particular corpus, hand, period, image quality, and scoring policy. Move any one of those and the number moves. This is not a rhetorical hedge; PRImA's survey of OCR evaluation tools and metrics found that even supposedly standardized CER and WER values cannot be compared directly between different implementations, because alignment, reading-order handling, and Unicode normalization differ between tools. The same survey reports that OCR confidence scores have at best a slight correlation with actual CER — so you cannot substitute a tool's own self-assessment for a measurement.
The published figures show the spread plainly. On nineteenth-century Fraktur novels, the OCR4all team reported an average CER of 0.85% for their pipeline against 5.3% for ABBYY FineReader — a specialist workflow on historical print, not a universal ranking. One multilingual benchmark reports figures around 1.7% CER on modern handwriting datasets, while the same benchmark's English historical subset puts Transkribus's Text Titan at 7.07% CER and 12.41% WER. Modern and historical material are effectively different tasks.
So the only figure that tells you what will happen to your wills, notarial acts, parish registers, or committee minutes is one you generate on your own pages. If you want the conceptual background on what these metrics do and don't measure before you start, read up on what AI transcription accuracy actually scores.
Step 1: Define the target before you look at any output (15 minutes)
Decide, in writing, what a correct transcription is for your project. This is the step most afternoon tests skip, and skipping it invalidates everything after it.
A diplomatic transcription preserves what is visibly on the page: abbreviations unexpanded, the long s as a long s, u/v and i/j as written, original spelling intact. A normalized transcription expands abbreviations and regularizes orthography for reading and searching. TEI's guidance on the representation of primary sources supports recording abbreviations alongside their intended sense, plus corrections, additions, deletions, substitutions, responsibility, and certainty — which is to say the distinction is not pedantry. It is the encoding decision your edition rests on.
These are different targets, not two versions of the same truth. If your reference expands a brevigraph while you penalize a tool for reproducing the visible mark, you are measuring your own inconsistency. Write down your policy on:
- Case
- Punctuation
- Whitespace and line breaks
- Unicode normalization and ligature handling (ct, st, æ, œ — one codepoint or decomposed?)
- Diacritics
- u/v and i/j
- Abbreviations and brevigraphs: retained or expanded
Hold that policy constant across every tool. OCR-D's ground truth guidelines treat these conventions as the foundation of any technically validatable reference, and the same discipline applies at a much smaller scale to an afternoon test. If you are unsure how to handle a particular mark, the practical decisions are laid out in this guide to manuscript abbreviations and ligatures.
Step 2: Sample randomly, then add a stress set (20 minutes)
Do not hand-pick your clearest pages. The temptation is strong and the result is worthless: a curated sample omits precisely the damage, marginalia, show-through, difficult hands, and awkward layouts that determine whether a tool is usable on the real collection. PRImA's survey notes that random sampling can support a statistically reliable statement where manual page selection cannot.
The random representativeness sample
Number the pages in your intended corpus and draw a small random sample — ten to twenty pages is realistic for an afternoon. Record the sampling frame and the draw. For context on scale, Europeana's historical-newspaper evaluation worked with 100 pages per content-holding partner and 100 pages per language, deliberately spanning digitization quality, formats, and scripts. Transkribus's own documentation suggests 15–30 pages for a single-hand collection and 50–100 across multiple hands; treat that as useful vendor planning guidance, not a field standard.
The stress set
Keep this separate and labelled: three to five pages you know are hard. The worst hand, the water-damaged leaf, the page with three marginal insertions, the one with a table. Score these separately and never average them into the random sample. Random sampling estimates average difficulty; the stress set tells you where the tool fails.
If your corpus spans multiple languages written in the Latin alphabet, sample across them and record the language of each page. Script and language are separate variables. A protocol built for English secretary hand applies equally to a French notarial act or a Dutch register, but the orthography, diacritics, and abbreviation system differ at each site, and averaging them into one number hides the variance. The distinction is developed further in this guide to which languages and scripts HTR can read.
Step 3: Freeze the inputs (10 minutes)
Every tool must receive identical pixels. This sounds obvious and is routinely violated — someone uploads a downsampled JPEG to one service and the original TIFF to another, or crops one set and not the other.
Fix and record:
- Resolution and file format
- Crop and rotation
- Colour or greyscale, and any binarization (ideally none — apply nothing you cannot apply identically everywhere)
- Model identity and version for every candidate, with the date
Version matters. A comparison run against "Claude" or "the Transkribus model" without a version string is not reproducible three months later. If your images came from a reading-room camera and legibility is variable, fix the capture stage before the test rather than after: uneven lighting and skew degrade every tool, but not proportionally. The practicalities are covered in this field method for photographing archival documents.
One more decision belongs here, and it is the one most afternoon tests get muddled. Are you asking "what can I use this afternoon?" or "could my own labelled pages produce a better local model?" Those are different experiments. Out-of-the-box and public models answer the first; fine-tuning answers the second. Mixing an untrained system with someone else's fine-tuned model answers neither fairly. For an afternoon, run the first experiment only, and label it as such.
Step 4: Key the reference (60–90 minutes)
This is the bulk of the afternoon, and there is no shortcut. Key each sampled page from the image, under your written convention, in a plain-text file linked to the image filename. Then proof it once against the image — ideally after a break, or by a second reader if you have one. OCR-D's guidance covers conventions and technical validation but does not establish a double-keying threshold, so treat an independent re-read as sensible quality control rather than a claimed standard.
One shortcut is legitimate: you can key the reference by correcting a machine transcription rather than typing from scratch, and this is common practice. The Humphries et al. study of eighteenth- and early-nineteenth-century English handwriting built its ground truth from a machine transcription that was then manually corrected and proofed. Be aware of the bias risk. If you correct output from Tool A and then score Tool A against it, you have anchored the reference to that tool's reading conventions. Either correct output from a tool you are not evaluating, or proof aggressively enough that every character has been independently checked against the image.
Keep the reference diplomatic if there is any chance expansion would erase evidence. You can always derive a normalized version from a diplomatic one; you cannot go back the other way.
Step 5: Run every candidate and score once (30 minutes)
Run the identical image set through each tool. Export plain text. Then apply your normalization policy to both sides — reference and output — with the same script or the same settings, before scoring.
CER is (substitutions + deletions + insertions) / number of reference characters. WER is the same edit operations over reference words. Concretely: reference `abcde`, output `abXde` is one substitution over five reference characters — 20% CER. Use one scorer for everything. JiWER is a Python package computing WER, CER, and related minimum-edit-distance measures, and is fine if you are comfortable with a script; dinglehopper compares ground-truth and OCR pages and reads ALTO, PAGE, and plain text. Transkribus's Compare view is convenient for non-programmers. Do not mix them: a score from one is not automatically comparable to a score from another unless tokenization, Unicode normalization, alignment, and reading order are fixed.
Retain per-page scores, not just the mean. Character errors are correlated within a page by hand, damage, and layout, so treating every character as independent overstates your precision. Per-page results let you see whether a mediocre average comes from uniform mediocrity or from two catastrophic pages.
If your pages have columns, marginalia, tables, or forms, report a second dimension. Vidal et al.'s page-level assessment of HTR argues for a two-fold evaluation: sequential WER, which is order-sensitive, alongside bag-of-words WER (bWER), which ignores order. The gap between them correlates with reading-order errors from layout analysis. A tool can recognize almost every word correctly and still serialize the page into unusable order — one number hides that entirely.
Step 6: Read the errors, don't just count them (30 minutes)
The scoring step ranks the tools. This step tells you which one you can trust. Go through the diffs and classify by hand:
- Substitutions, insertions, deletions at character level — the ordinary, visible kind of recognition error.
- Abbreviation and diacritic errors — was the macron dropped? the brevigraph silently resolved? the accent normalized away?
- Layout and reading-order failures — columns interleaved, marginalia inserted mid-sentence, table cells flattened.
- Silent normalization — archaic spelling "corrected," u/v regularized, the long s read as f or quietly modernized.
- Fluent fabrication — plausible text that is not on the page. A name that reads like a name. A date that reads like a date. A legal formula completed from the model's expectation rather than from the ink.
That last category is worth an afternoon on its own. A garbled character error announces itself; you see it, you check the image, you fix it. A fluent invention passes a casual proofread and propagates into your footnotes. State the evidence carefully: studies document language imbalance, degraded performance on historical and non-English material, and cases where automatic post-correction reduced accuracy, but the literature does not settle how much of a fluent wrong reading comes from image downsampling, language-model priors, or prompt context. Treat fluent-but-wrong as a documented risk you test line by line rather than a law of nature — and note that a plausible-sounding transcription is precisely the kind you are least likely to catch. The mechanics of this failure mode are worth understanding before you evaluate any general model: see fluent but wrong.
One honest note on the published comparisons. Humphries et al. reported that on their eighteenth- and early-nineteenth-century English corpus, Claude Sonnet-3.5 scored a strict CER of 7.3% against Transkribus Titan's 8.0% — a result on a single English corpus, with the paper itself noting the LLM results remained roughly 60% less accurate than the upper error bound for non-expert human transcribers, and with no equivalent evidence for Dutch ledgers, Spanish or Italian archival hands, or heavily abbreviated legal material. It is a preprint, one corpus, one language, and everything still requires verification against the image. It is a reason to include a general model as a control in your test, not a reason to skip the test.
What to record so the result is defensible
Write up, in a paragraph you can paste into a methods note:
- Target (diplomatic or normalized) and the full normalization policy
- Sampling frame, sample size, and how the draw was made; stress set listed separately
- Image resolution, format, crop, and any preprocessing
- Every tool with model name and version, training status, and the date of the run
- Scorer used
- Per-page CER and WER (plus bWER if layout matters)
- Qualitative error categories with examples
Call the result what it is: a directional local benchmark for this corpus on this date. That framing is stronger, not weaker — it is a claim you can defend in a footnote.
Where Leo fits in a test like this
Leo is one of the candidates you would put in the run, and it is worth being direct about why we think it holds up on this kind of material.
The relevant claim is the one measured under exactly this protocol. On a randomized 97-image sample of early-modern English manuscripts from the Folger Shakespeare Library, at ATR-1's release, Leo scored approximately 5% CER — 61% fewer errors than the next-best model, with Transkribus (Text Titan I) at ~13%, Claude Opus at ~23.3%, Gemini 2.5 Pro at ~24.8%, and GPT-4.1 at ~56.7%. The full results are published here. That is one corpus, one script, one period — the same caveat we apply to everyone else's numbers applies to ours, which is the whole argument for running your own.
Three things make Leo straightforward to include in an afternoon test. ATR-1 is zero-shot: there is no model to train or select before you get output, so it answers the "what can I use this afternoon?" question directly rather than requiring a separate fine-tuning experiment. It reads Latin-script material — any language written in the Latin alphabet, handwritten or printed — so a mixed-language sample can go through in one pass. And it is built to transcribe what is on the page rather than smooth it: strikethroughs, additions, marginal notes, expansions, and archaic orthography survive, which means your diplomatic reference and Leo's output are measuring the same target rather than talking past each other.
Two honest boundaries for your test design. Complex tabular layouts vary in quality, and the known weak spot is pages where dominant printed structure meets dense handwriting — pre-printed ledger and deed-book forms, where the model can favour the printed headers over the manuscript entries. If your corpus is mostly that, put it in your stress set and look hard at the result. Separately, if you want a general model as a control, you can run GPT, Gemini, or Claude as additional transcription tabs inside Leo on the same images, which at least removes input variance from that comparison.
After the test
A tool comparison is not the end of the workflow. It is a gate into it. Whichever tool you choose, the transcription that comes out is a first pass, and the discipline that follows — checking names, dates, sums, and legal formulae against the image, keeping editorial intervention in a layer separate from the base reading, recording what you changed and why — is the part that makes the work citable. A working method for verifying transcription accuracy picks up where this protocol leaves off, and the broader question of how machine reading fits into a scholar's practice sits in the wider field of handwritten text recognition.
The afternoon is well spent for a reason that outlasts the tool choice. Keying twenty pages of reference under a written convention, then reading a machine's errors against them line by line, is paleographical training with immediate returns. You will finish knowing which letterforms in your particular hand are genuinely ambiguous, which abbreviations your corpus uses habitually, and which of your own readings you had been guessing at. That knowledge stays with you when the models change — and they will change. The protocol is the durable asset.
Frequently Asked Questions
How do I test AI transcription on my own documents?
Run a six-step protocol in a single afternoon. Write down your transcription target first — diplomatic or normalized — along with your policy on case, punctuation, whitespace, Unicode normalization, diacritics, u/v and i/j, and abbreviations. Draw a small random sample from the collection you actually intend to transcribe, and keep a separate stress set of three to five known-hard pages. Freeze the images so every tool receives identical pixels. Key a careful reference transcription. Run each candidate, score character error rate and word error rate with one scorer under one normalization policy, then read the errors by hand and classify them. The result is a directional local benchmark for that corpus on that date.
How many pages do I need to test transcription accuracy?
Ten to twenty randomly sampled pages are realistic for an afternoon and give you a directional local benchmark, not a collection-wide figure with narrow uncertainty. Number the pages in your intended corpus, record the sampling frame and the draw, and resist hand-picking your clearest pages — a curated sample omits the damage, marginalia, show-through, and awkward layouts that decide whether a tool is usable. Keep three to five deliberately difficult pages as a separate stress set, scored on their own and never averaged in. Transkribus's own documentation suggests 15–30 pages for a single hand and 50–100 across multiple hands; treat that as vendor planning guidance rather than a field standard.
What is the difference between CER and WER in transcription testing?
CER is character error rate: substitutions plus deletions plus insertions, divided by the number of reference characters. If your reference reads `abcde` and the output reads `abXde`, that is one substitution over five characters — 20% CER. WER applies the same edit operations to words instead of characters. Use one scorer for every tool, because CER and WER values are not directly comparable between implementations: alignment, reading-order handling, tokenization, and Unicode normalization differ. Keep per-page scores rather than only the mean, since errors cluster by hand, damage, and layout. If your pages have columns, marginalia, or tables, also report order-insensitive bag-of-words WER alongside sequential WER.
Can I build my reference transcription by correcting machine output?
Yes, and it is common practice — the Humphries et al. study of eighteenth- and early-nineteenth-century English handwriting built its ground truth from a machine transcription that was then manually corrected and proofed. The bias risk is specific: if you correct output from Tool A and then score Tool A against that reference, you have anchored the reference to that tool's reading conventions. Either correct output from a tool you are not evaluating, or proof aggressively enough that every character has been checked independently against the image. Keep the reference diplomatic where expansion could erase evidence; you can derive a normalized version later, but not the reverse.
Why are fluent transcription errors more dangerous than garbled ones?
A garbled character error announces itself — you see it, check the image, and fix it. A fluent invention reads like a name where a name belongs, or a date where a date belongs, or a legal formula completed from the model's expectation rather than from the ink. It passes a casual proofread and propagates into footnotes. Studies document language imbalance, weaker performance on historical and non-English material, and cases where automatic post-correction reduced accuracy, but the literature does not settle how much fluent misreading comes from downsampling, language-model priors, or prompt context. Treat it as a documented risk you test line by line, not a law of nature.