Writing a Digitization Services RFP: Transcription Requirements You Can Actually Test
How to draft enforceable transcription-accuracy clauses in a digitization services RFP, using measurable CER, buyer-owned ground truth, normalization rules, and separate image and text acceptance gates.
Leo Team
August 6, 2026

Contents
This is a working guide to the text-accuracy clauses in a digitization services RFP — the ones that decide whether you can prove, at delivery, that the transcription you paid for is sound. Image capture already has named standards and star levels; transcription usually gets a vague adjective and a vendor's headline percentage. What follows is how to close that gap with terms you can measure and enforce.
A digitization services RFP is only as strong as the requirements you can measure at delivery. Most fail on the text side: image quality gets a named standard and a star level, while transcription gets "high accuracy" and a number from a vendor's marketing page. The fix is structural. Specify image capture against FADGI, ISO 19264-1, or Metamorfoze; specify text accuracy against a character error rate measured on a buyer-selected, stratified, blind sample of your own series, with normalization rules written down in advance; and make image acceptance and text acceptance two separate gates that a vendor must pass independently.
The sections below cover the clauses that hold up under scrutiny — and the ones that quietly transfer risk back to your office.
Why "accurate transcription" is not a requirement
Write "the supplier shall deliver highly accurate transcriptions of all records" into a solicitation and you have written nothing enforceable. Accurate against what reference? Measured how? On which pages? With what treatment of abbreviations, capitalization, and archaic spelling? When the first batch arrives and someone in your office says the names look wrong, you have no instrument to prove it and no clause to invoke.
The measurable version is character error rate. OCR-D defines CER from Levenshtein edit distance: substitutions plus deletions plus insertions, divided by the number of characters in the reference transcription, expressed as a percentage. Word error rate applies the same arithmetic after tokenizing into words. Both are simple to compute once two things exist: a reference transcription (ground truth) and a written normalization policy.
CER is generally the more stable measure for historical and vernacular handwriting, because word boundaries and spelling conventions are not stable in the sources themselves. Hodel and Schoch, in their study of general models for German Kurrent, used CER precisely because they had no conclusive definition of a "word" in their material. Court registers, deed books, and probate files present the same problem: inconsistent spacing, abbreviations that may or may not be one token, clerks who join and split words at will.
CER has a real weakness for record series, and the RFP has to compensate for it. An aggregate CER averages across thousands of characters and can look excellent while hiding one wrong digit in a case number, a parcel identifier, a date, or a surname. That is why the text-accuracy clause needs two tiers: an overall CER threshold, and separate exact-match or per-digit accuracy requirements on the fields your users actually search.
The ground truth clause
Ground truth is an image-linked reference transcription against which a system's output is scored. If the vendor supplies it, you are not testing the vendor. Write the production of ground truth into the RFP as a buyer-controlled activity, or as an activity performed by an independent reviewer under buyer instruction.
A defensible ground-truth set has five properties, and each is a sentence you can put in a solicitation:
- Buyer-owned random selection. The buyer or independent reviewer draws the sample from the actual series. Record the sampling frame and the random seed in the acceptance report.
- Independent double-keying or independent verification. Two keyers work separately; a third adjudicates disagreements.
- A documented transcription policy. Written before keying, not negotiated after the results come in.
- Adjudication of disagreements, with the reasoning recorded.
- Preservation of unresolved loci. Where two competent readers cannot agree what the clerk wrote, that uncertainty is part of the record, not something to be settled by fiat so the arithmetic is tidy.
OCR-D's Ground Truth Guidelines exist to make ground truth technically validatable and to allow existing transcriptions to be checked against a common basis. Citing them by name in the RFP saves you drafting a transcription policy from nothing.
Normalization: the clause that decides the score
Before any score is meaningful, the RFP must predefine how the comparison treats:
- case and whitespace
- punctuation and diacritics
- Unicode normalization form
- abbreviations and brevigraphs — is a preserved contraction mark correct, or is the expanded form correct?
- editorial expansions — `yo[u]r` scored as written, or normalized before comparison?
- tokenization, if WER is reported at all
None of this is decoration. Whether a system that silently expands abbreviations scores well or badly depends entirely on which of these rules you wrote. Normalization remains genuinely contestable for abbreviations, spelling variation, supplied readings, and uncertain characters. You do not need the universally correct policy. You need a policy, stated in the RFP, applied identically to every proposal.
For judicial and land records, the policy question that matters most is whether you want the record as written or the record as modernized. For anything that may be relied upon legally, the answer is as written — the transcription should reproduce what is on the page, with interpretation kept in a separate layer. A guide to why source integrity is the constraint that sets everything downstream is worth reading before you draft this paragraph.
The acceptance sample
A fair acceptance sample is randomized from the actual series and stratified by the variables that drive difficulty: language, script and hand, date or period, document type, layout, image quality and condition, and any known problem areas. A sample drawn from your cleanest 1870s docket volumes predicts nothing about the 1790s loose papers.
Keep any development or tuning data the vendor sees strictly separate from a blind acceptance holdout. Report results overall and by stratum, with confidence intervals. A contract that accepts on aggregate CER alone will accept a system that reads your clean material well and your worst hands not at all.
On sample size, be honest with yourself: there is no standard answer. NIST's acceptance-sampling guidance supports sampling precisely where 100% inspection is prohibitively expensive, and a NIST evaluation of stratified randomized acceptance sampling works through sample sizes of n = 50, 75, 100, 125, 150 and 175. Those are study designs, not a digitization rule. Correct sample size depends on lot size, your confidence level, allowable error rate, and stratum minimums — which is to say, on consequences. A series that supports property title carries different consequences from a series of nineteenth-century administrative correspondence, and the RFP should say so.
Setting a threshold you can defend
The central gap in procurement is this: no recognized cross-domain standard comparable to FADGI or ISO 19264 sets a universal CER or WER threshold for historical transcription. Image quality is a solved specification problem. Text accuracy is not.
The published figures show why a borrowed number is dangerous. Hodel and Schoch report validation CERs ranging roughly from 2.55% to 7.80% across their German-Kurrent corpora — Early Modern Kurrent in a single hand, medieval charters in three or four hands, nineteenth-century Zürich state-archive material in about twelve — and note that some validation sets share hands with training. The Folger Shakespeare Library's experiments with early modern manuscripts report CER falling from 37% to 21% once models were trained on historical data — a large project-level improvement, and still nowhere near a threshold anyone would write into a contract. Work on TrOCR ensembles for sixteenth-century Latin-language manuscripts reports ensemble CER around 1.60 against a non-augmented baseline of 1.93. That is a different corpus, a different language, a different degree of regularity. It tells you nothing about a Dutch notarial register or an Alabama chancery docket.
No independently corroborated, apples-to-apples CER table spanning English, French, Spanish, Dutch, German and Italian government series appears to exist. So a vendor's headline number, whatever its provenance, is a capability statement about someone else's material.
The procedural answer is a paid pilot before the threshold is fixed. Structure the solicitation in two stages: a scored pilot on a buyer-selected sample of your own series, with results reported by stratum, and a production contract whose acceptance threshold is set from those pilot results plus your risk weighting. Publish the pilot sample composition in the RFP so bidders are pricing the same work. This is the single change that converts a text-accuracy clause from aspiration into an enforceable term.
Two things that look like accuracy and are not
Confidence scores
A confidence score is a model's statistical estimate that a detected result is correct — Microsoft's documentation states this plainly — not an observed proportion correct against gold transcription. A vendor reporting "average confidence 0.94" has reported nothing about accuracy. Require CER, WER where tokenization is defined, and critical-field accuracy measured against independent ground truth. Confidence scores may still earn a place in the contract as a routing mechanism — flagging pages for human review — but never as an acceptance metric.
Image-standard compliance
FADGI's Third Edition (May 2023) sets four star-level image-quality goals. ISO 19264-1 specifies an imaging-system quality analysis method using defined targets — name the adopted edition in the contract, since a draft second edition exists. Metamorfoze Image Quality v2.0 (April 2025) specifies measurable relationships between preservation images and originals. NARA's Digitization Quality Management Guide covers capture-device performance testing and inspection, and NDSA's Levels of Digital Preservation let you assess preservation capability. All of it is testable. None of it says anything about whether the text is right. Two gates, scored separately.
Deliverable formats: specify the package, not "searchable PDF"
Searchable PDF preserves a visual page with embedded text. It is a reasonable access derivative and a poor sole transcription deliverable, because the text is bound to a presentation format rather than exposed for reuse. Plain UTF-8 is portable but discards page, line, and word coordinates, reading order, editorial markup, and provenance unless sidecar files supply them.
Specify the layers you actually need:
- ALTO XML or PAGE XML for analyzed layout and text objects: page, region, line and word positioning, reading order, recognition attributes. This is what makes word-level highlighting over an image possible in a public access system.
- TEI for semantic and editorial transcription — unclear text, supplied readings, abbreviations, revisions, source description. TEI is not primarily a layout schema; do not ask it to do ALTO's job.
- METS to wrap descriptive, administrative and structural metadata and bind images to their ALTO/PAGE/TEI and PREMIS records. METS is a container, not the transcription content model.
- PREMIS for preservation objects, events, agents and rights.
- BagIt (RFC 8493) for transfer packaging with manifests, plus logged fixity checks so that chain of custody is demonstrable between supplier and repository.
If you are still deciding which of these layers your access goals require, settle the distinction between finding-aid description and full-text transcription first. It determines most of the rest of the pipeline you are procuring, and how the resulting text will need to be described once it exists.
Acceptance language: delivery is accepted only when schemas validate, checksums and manifests verify, every text unit links to its source image, and provenance events including model name and version are supplied.
Clauses specific to judicial and government records
Court and registry series carry obligations that heritage digitization contracts often omit.
Rights
Paying for scanning does not automatically give you the text. Ownership and reuse depend on the contract, applicable law, and any pre-existing vendor material embedded in the workflow. Assign or license expressly: images, recognized text, human corrections, annotations, layout data, metadata, ground truth, processing logs, and derivatives. Government buyers with a federal nexus should note how DFARS Part 227 handles rights in technical data and copyright licenses; NARA's non-exclusive digitization agreement is a useful model for partnership arrangements. If a vendor proposes retaining your text or using it as training data, that is a rights term, and it belongs in the rights clause rather than buried in a data-processing appendix.
Restricted content
Redaction correctness, sealed-record handling, and no unauthorized disclosure should be enforceable deliverables, not process promises. Federal court practice on redaction requirements and sealed documents and Florida's Standards for Access to Electronic Court Records are worth citing where they apply. European and UK archives operating under data-protection law should reconcile the archiving-in-the-public-interest position set out in the National Archives' GDPR FAQs with the vendor's processing terms before signature.
Process and service levels
Trained handlers, fragile-original handling procedures, secure transport, role-based access, incident response, subcontractor controls, retention limits, batch turnaround, and return-or-erasure certification at contract close. Collier County's court-records digitization RFP is a serviceable public example of how these are phrased in a live solicitation.
Human review as a scoped line item
If the series contains high-stakes fields, require a defined review tier rather than leaving it implicit. The same logic applies whether the labor is a vendor's staff or a volunteer programme: Transcribe Bentham's frequently quoted 94% staff-approval rate is a workflow statistic, not a CER, and Causer and colleagues are explicit about what it does and does not measure. Crowdsourced or vendor-keyed, human output needs the same sampling, adjudication, and acceptance metrics as machine output. A hybrid pipeline with targeted review of names, dates, and identifiers is usually the defensible design for judicial series.
Testing bidders on your own series before the contract is signed
The most useful thing you can do during evaluation is score every bidder on the same buyer-selected sample. That includes scoring the software you might run in-house, since for many series the realistic choice is not "vendor or nothing" but "outsourced service or an internal team with a transcription platform."
That evaluation is where Leo is worth putting on the list. Its transcription model, ATR-1, is zero-shot: it reads Latin-script material out of the box, with no per-series model training step, which matters when your pilot has to be scored in weeks rather than after a training cycle. Scope is defined by the alphabet, not the language — English dockets, French notarial records, Dutch registers, German parish books and Spanish land papers are all in scope; Greek, Cyrillic, Hebrew, Arabic, and Indic or East Asian scripts are not. Printed matter of any period is in scope alongside manuscript hands. And it is built to transcribe what is on the page rather than smooth it into modern prose — which is the behaviour your normalization policy is trying to detect in the first place, and the reason a fluent-sounding general chatbot transcript is the hardest kind of output to audit.
Where head-to-head accuracy is the question, here is the one published comparison: on a randomized 97-image sample of early-modern English manuscripts from the Folger Shakespeare Library, at ATR-1's release, Leo scored roughly 5% CER against Transkribus/Text Titan I at about 13%, Claude Opus at 23.3%, Gemini 2.5 Pro at 24.8% and GPT-4.1 at 56.7% — 61% fewer errors than the next-best model, with the full comparison published here. Apply the same scepticism you would apply to any bidder's number: that is early-modern English manuscript material, not your series, and the point of your pilot is to find out what happens on yours.
Two boundaries to price in honestly. Leo produces UTF-8 text with TEI XML as its scholarly export; it does not emit ALTO or PAGE, and does not handle IIIF, OAIS packaging, or EAD/METS/PREMIS — those layers stay with your repository system or your capture vendor. And pages dominated by pre-printed forms with dense handwriting filled in — some ledger and deed-book layouts — are its weaker material, which is exactly the kind of thing a stratified pilot sample should be designed to expose rather than avoid. Free credits are available for archives through the Leo Transcription Grant, on the condition that the resulting transcriptions and images are published openly, which suits some public series and not others.
What a good RFP actually buys you
The clauses above do not guarantee good text. They guarantee that you will know whether the text is good, at the moment when you still have leverage — before final payment, before ingest, before a genealogist or a title examiner relies on it.
That is the real deliverable of a well-drafted solicitation: not a promise, but an instrument. A named image standard with a star level. A written normalization policy. A blind, stratified sample you selected. A threshold derived from a pilot on your own material rather than from someone else's corpus. Per-field accuracy on the identifiers your users search. A package that validates and a chain of custody that verifies.
Write those and you have converted the vaguest part of a digitization contract into something you can enforce, and something a successor sitting in your chair in ten years can audit. The records were kept carefully once. The specification is how that care survives the transfer into machine-readable form.
Frequently Asked Questions
What should a digitization services RFP include for transcription accuracy?
A digitization services RFP should specify text accuracy as a character error rate measured on a buyer-selected, stratified, blind sample drawn from your own series, with a written normalization policy fixed in advance. Add separate exact-match or per-digit requirements for the fields users actually search — case numbers, parcel identifiers, dates, surnames — because an aggregate CER can hide a single wrong digit. Keep image acceptance and text acceptance as two independent gates: name FADGI, ISO 19264-1 or Metamorfoze for capture, and score transcription separately against independent ground truth.
What is an acceptable character error rate for historical handwriting?
There is no recognized cross-domain standard setting a universal CER or WER threshold for historical transcription, unlike image quality where FADGI and ISO 19264 apply. Published figures vary widely by corpus: validation CERs from roughly 2.55% to 7.80% across German Kurrent collections, a fall from 37% to 21% on early modern English manuscripts once models were trained on historical data, and around 1.60 for ensembles on sixteenth-century Latin. None of those transfer to your series. Set the threshold from a paid, scored pilot on your own material, reported by stratum, plus your risk weighting.
Can a vendor's confidence score be used as an acceptance metric?
No. A confidence score is a model's statistical estimate that a detected result is correct, not an observed proportion correct measured against gold transcription. "Average confidence 0.94" says nothing about accuracy. Acceptance should rest on character error rate, word error rate where tokenization is defined, and critical-field accuracy scored against independent ground truth that the buyer or an independent reviewer produced. Confidence scores can still earn a place in the contract as a routing mechanism — flagging pages for human review — but they belong nowhere near the acceptance clause.
What deliverable formats should a digitization contract specify instead of "searchable PDF"?
Specify the layers your access goals require rather than a single format. Searchable PDF is a reasonable access derivative but a poor sole transcription deliverable, since the text is bound to a presentation format. ALTO XML or PAGE XML carry page, region, line and word coordinates, reading order and recognition attributes — the basis for word-level highlighting over an image. TEI handles semantic and editorial transcription. METS wraps descriptive, administrative and structural metadata; PREMIS covers preservation events and rights; BagIt handles transfer packaging with manifests and logged fixity checks.
Who owns the transcribed text after a digitization vendor delivers it?
Ownership is not automatic — paying for scanning does not by itself give you the text. Rights depend on the contract, applicable law, and any pre-existing vendor material embedded in the workflow. Assign or license expressly: images, recognized text, human corrections, annotations, layout data, metadata, ground truth, processing logs, and derivatives. If a vendor proposes retaining your text or using it as training data, treat that as a rights term and place it in the rights clause rather than a data-processing appendix. Government buyers with a federal nexus should also check how DFARS Part 227 handles technical data and copyright licences.