Planning a Mass Digitization Project: From Collection Survey to Searchable Text

Planning a mass digitization project stage by stage—from survey and prioritization through capture, text generation, QA, and publication—with focus on the text layer and verification gates.

Leo Team

August 18, 2026

Planning a Mass Digitization Project: From Collection Survey to Searchable Text
Contents

A mass digitization project runs in a fixed order, and each stage constrains the one after it. This guide walks the sequence — survey, rights, prioritization, capture, text generation, description, QA, publication, preservation — with particular attention to the two stages institutions consistently under-plan: generating the text layer, and proving it worked. It is written for archivists, librarians, and museum staff planning at collection scale.

A mass digitization project is a coordinated programme, not a scanning campaign. The decisions that determine whether a collection ends up genuinely searchable are made at the survey stage, months before the first image is taken, because the intended search outcome dictates capture specification, text-generation route, and QA sampling. The most common planning failure is treating the scanner as the finish line: a page image is not searchable text, and, as the Northeast Document Conservation Center puts it plainly, "digitization is not preservation — it is simply a means of copying original materials."

The scope below covers a few thousand pages or a few million, and it sits within the broader question of how archives digitize, describe, and make collections findable.

Start from the search outcome, not the shelf

Before the survey, answer one question: what will a user type, and what should come back?

Three answers are common, and they are not the same project.

Finding-aid discovery

A researcher searches for a collection, series, or box and gets a description. This needs good descriptive metadata and images. It does not need full text.

Retrieval-grade full text

A researcher searches a surname, a place, a ship's name, and gets a hit list of pages. This needs a text layer per page, indexed and linked to the image, and it tolerates some error — a name misread once may still be found via a variant or a fuzzy match.

Scholarly or diplomatic transcription

A researcher cites the text as evidence. This needs verbatim fidelity to the page, encoded conventions, and human verification of the passages that carry weight.

Most institutions want the second and are quietly budgeting for the first. The distinction between metadata-level and full-text discovery is worth settling in writing before anything else, and it maps onto specific pipeline choices worked through in more detail in planning archive searchability.

Stage one: Survey — should, may, can

The survey is a triage exercise with three separate tests, and NEDCC's framing remains the most useful: decide whether the material should be digitized, whether it may be digitized, and whether it can be. There are, as NEDCC notes, "no absolute criteria" for selection — only questions answered in the context of your own institution.

Record, per series:

  • Extent. Page counts, not linear feet. Every downstream cost line is per-image.
  • Condition. Iron-gall ink burn, foxing, brittle margins, tight bindings, show-through. Anything that will need conservation before capture, or a different capture rig.
  • Format and layout. Bound volumes versus loose sheets; single-column prose versus ledger grids versus pre-printed forms with manuscript entries; oversized maps and plans.
  • Script and language. Which hands, which languages, which centuries. Note where a series changes clerk mid-run, because recognition quality often changes with it.
  • Rights. Whether you hold or can clear the right to publish images and text. Unpublished archival material can carry third-party copyright, and this is the constraint most likely to strand a finished dataset.
  • Demand. Reading-room request logs, reference-enquiry patterns, citation trails. Demand is the cheapest prioritization signal you have.

Two survey fields do disproportionate work later. The first is layout complexity, because mixed printed-and-handwritten pages and dense tabular grids are the hardest material for any recognition system and should be routed and evaluated separately from prose. The second is hand consistency, because a series in one clerk's hand behaves very differently from a series that turns over every three years.

Stage two: Prioritization

Weight the survey. A workable scheme scores each series on demand, fragility (higher fragility argues for digitization to reduce handling), rights clarity, and text-generation difficulty, then sequences accordingly.

A practical sequencing rule: put one easy series and one hard series in the first phase. The easy series builds throughput and trains staff. The hard series surfaces the problems — the ledger forms, the faded ink, the layout that no one anticipated — while there is still time to change the plan. Institutions that front-load only easy material discover their real constraints in year three.

Stage three: Capture

Capture is the best-specified stage in the whole pipeline, which is why it is also the stage that lulls planners into confidence. Imaging quality is governed by measurable guidance: FADGI's Technical Guidelines for Digitizing Cultural Heritage Materials, Third Edition, developed by the Still Image Working Group in 2022–2023, sets shared best practice for still-image materials and is explicitly designed to be used with conformance evaluation targets and testing software — the guidelines plus the monitoring, together, are what make a programme FADGI-conforming. ISO 19264-1 describes a method for analysing imaging-system quality specifically in cultural-heritage imaging. Metamorfoze remains the third reference point for preservation-grade imaging in European practice.

Three planning notes.

Capture is roughly a third of the budget

NEDCC reports, cautiously, that high-resolution digital capture consumes "perhaps one-third of the cost of a digitization project." That figure is experience-based rather than a current per-page price, but as a sanity check it is useful: if capture is the whole of your budget, you have not budgeted for the project. The remaining two-thirds is metadata, text, QA, publication, and storage. A fuller line-item view is set out in what a records digitization project actually costs.

Specify capture for the text you intend to generate

Recognition models work from pixels. Resolution, evenness of lighting, and geometry all affect legibility of fine diacritics, superscript abbreviation marks, and thin ascenders. If the plan includes full text, do not economise on capture and expect recognition to compensate. The stage-by-stage relationship between capture standards and everything downstream is worked through in building an archival digitization workflow.

Handle damaged material as a routing decision, not a hope

Faded iron-gall ink, foxing, and show-through are three different problems with three different remedies — some are capture problems, some recognition problems, some conservation problems. Sorting them before capture saves rework; the triage logic for damaged documents is worth settling at this stage.

Stage four: Text generation — the stage that decides the project

This is where mass digitization projects succeed or quietly fail. The image gives you a surrogate. The text layer is what makes the collection findable, and it has to be generated by a separate process: OCR, handwritten text recognition (HTR), human transcription, or some combination.

Route by material, not by preference.

Modern and clean print

General-purpose OCR is the sensible first route. It is fast, cheap, and accurate on high-contrast modern type. Benchmark it on a sample and move on.

Historical print

Do not assume the same tools carry over. The Early Modern OCR Project documents why: hand-press printing, roughly 1475–1800, produced texts with fluctuating baselines, mixed founts, and varied concentrations of ink, and much of it survives in poor-quality images. eMOP's response was typeface analysis, engine training, page-image comparison tooling, and crowd-sourced correction — a customization workflow, not a switch to flip. Conventional OCR is engineered around regular letterforms and predictable spacing; historical typography violates the assumptions. The long s comes back as f. Ligatures split or drop. The macron standing for an omitted nasal has no representation in a glyph classifier, so it is dropped or silently normalized. Archaic spelling and u/v interchange get "corrected" by post-processing that treats period orthography as error.

Handwriting

Handwriting needs HTR or human transcription. OCR does not read manuscript hands reliably, and planning as though it does is the single most expensive misconception in this work. Within HTR, the practical planning question is trained versus ready-to-use. Transkribus's central strength is training custom handwritten text recognition models, and eScriptorium and OCR4all work the same way. That path can produce strong results on a consistent series — but it front-loads ground-truth labour, and the peer-reviewed OCR4all paper is candid that manually finding and transcribing text lines for training data is a highly non-trivial task. Budget it as staff time with a real page count, or choose a route that does not require it. Whether your series actually justifies the training investment is worth deciding deliberately rather than by default — the case for testing a ready model first applies to most institutional collections.

Crowdsourcing

Volunteer transcription remains genuinely valuable, particularly for exact verbatim reading, correction of machine output, and ground-truth creation. The Smithsonian Transcription Center asks volunteers to record the exact letters and words on the page, which is precisely the discipline a scholarly transcription needs. But crowdsourcing is not free and not infinitely elastic. The Transcribe Bentham evaluation studied over 4,000 volunteer-submitted transcripts checked by staff across a 20-month period, and found staff checking and approval averaged 3.5 minutes per volunteer page — an efficient rate, but staff time nonetheless, and the same project evaluation estimated that at the then-current rate Bentham's papers would be complete by 2036. That is a forecast about one exceptionally well-supported project, not a scaling law. Plan crowdsourcing series by series, with review capacity costed — the division of labour between volunteers and machine first-pass is a per-series judgment.

What not to route to a general chatbot

Uploading page images to a general-purpose assistant is the fastest way to get text that reads well and cannot be trusted. General models downsample images and lean on linguistic plausibility, so they fill gaps with fluent invention rather than leaving them visibly broken. Garbled OCR announces its failures. A confident paraphrase does not, and at collection scale nobody re-reads 40,000 pages to find it. The mechanics of this failure mode are set out in fluent but wrong.

Where a ready model fits in the plan

For collections where the material is Latin-script — the alphabet, not the language, so English wills, French notarial registers, Dutch church books, Spanish and Italian correspondence, German records, and Latin-language documents all sit inside the same scope — a zero-shot specialist model changes the shape of the text-generation stage, because it removes the ground-truth phase from the critical path. Leo's ATR-1 is built for exactly this: handwritten and printed Latin-alphabet material from roughly the past five hundred years, including secretary hand, Gothic cursive, and italic, with no per-collection training, no segmentation, and no preprocessing step to specify in the project plan. On a randomized 97-image sample of early-modern English manuscripts from the Folger Shakespeare Library, at ATR-1's release, Leo recorded approximately 5% character error rate against roughly 13% for Transkribus's Text Titan I and 23–25% for Claude Opus and Gemini 2.5 Pro — the full comparison is published here. One corpus, one language, one release; treat it as directional and benchmark on your own pages.

Two things matter more than the number for a digitization plan. The first is source integrity: ATR-1 transcribes what is on the page — strikethroughs, interlinear additions, marginalia, expansions, archaic spelling, the long s, the abbreviation mark left unresolved rather than silently expanded. That is what makes machine output usable as the base layer of an archival record rather than a paraphrase of one. The second is that the analysis stays separate. Transformations — Correct, Modernize, Translate, Summarize, Classify, Extract named entities — each write to a new tab and never overwrite the base transcription, so an interpretive layer never contaminates the record. Translation is worth stating explicitly because it is the most common confusion: transcription and translation are two operations, and Leo does both, separately.

Around the model sits the workflow a digitization programme actually needs day to day: batch upload including multi-page PDFs, nested folders, per-document archival metadata (Title, Creator, Date, Archive, Type, Collection, Box, Folder, Identifier, Rights), fuzzy full-text search across everything transcribed, public read-only share links, and export to TEI XML, Word, PDF, or HTML. Leo begins at upload — it does not capture — and it does not produce IIIF manifests, OAIS packages, or EAD/METS/PREMIS, which your repository and delivery layer handle. Two honest limits worth recording in a plan: pre-printed ledger and deed-book forms with dense manuscript entries are the known weak spot, where the model can favour printed headers over the written entries, and complex tabular layouts vary. Route those series into their own sample and evaluate them on their own terms. Non-Latin scripts — Greek, Cyrillic, Hebrew, Arabic, Indic, East Asian — are out of scope entirely.

Academics, students, and archives can apply for up to 100,000 free credits through the Leo Transcription Grant, on the condition that the resulting transcriptions and images are published openly within 24 months — which means applying with material you hold or can clear the rights to publish. Printed material is not eligible for the grant; for print-heavy programmes, the free tier or a paid plan is the route.

Stage five: Description, encoding, and keeping text tied to image

Two rules govern this stage.

Keep the image and its text linked, permanently

Every machine transcription is a claim about a page, and the claim is only checkable if the page is one click away. This is also the operational basis of all verification: reviewers read the text against the image, never in isolation.

Layer your standards rather than overloading one

Dublin Core, via DCMI Metadata Terms, handles general discovery metadata. Richer archival, structural, and preservation description belongs in EAD, METS, MODS, and PREMIS, catalogued among the Library of Congress standards. For the text itself, TEI is the scholarly encoding choice, while ALTO carries technical OCR metadata including word-level coordinates — the thing that lets a search hit highlight on the image. Image delivery and reuse is IIIF's job, conceived to enable systematic reuse of image resources across cultural-heritage repositories; it delivers images, not text, so full text and metadata are layered alongside it. Which schema carries what is set out in more depth in archival metadata standards for transcribed documents.

Stage six: Quality assurance that can fail

QA that cannot fail is not QA. Build two independent acceptance gates.

Image gate

Conformance against your chosen imaging target, tested with software, on a defined sample. FADGI's model — guidelines plus targets plus monitoring — is the pattern.

Text gate

Character error rate and word error rate against buyer-created ground truth, measured on a random sample of representative pages, per series. CER and WER are the established recognition-evaluation measures, and the OCR-quality assessment literature confirms their use in exactly the mass-digitization context. What that literature does not supply — and what no source in current practice supplies — is a validated universal threshold for "good enough for search." Anyone who quotes you one is guessing. Set your own, tie it to your intended use, and weight the sample toward the tokens that matter: names, dates, places, sums, boundary calls. A 4% CER spread evenly across function words is a different collection from a 4% CER concentrated in surnames. The method for verifying transcription accuracy is worth writing into the project documentation, and if you are procuring transcription externally, accuracy clauses have to be testable to be enforceable.

Sample the hard material separately. An aggregate error rate averaged across clean prose and ledger grids tells you nothing actionable about either.

Stage seven: Publication and preservation

Publication means the search index, the delivery interface, the rights statements, and the link from hit to page. Preservation means master files in managed storage with fixity checking, format documentation, and a plan for migration. The two are separate obligations and separately funded, and the surrogate does not discharge the preservation duty for the original.

Then plan the next phase from what the last one measured: throughput per series, actual cost per page including QA, error rates by material type, and the routing decisions that turned out to be wrong. A digitization programme that does not feed its own numbers back into its own planning is running on the estimates it made before it knew anything.

The part no model plans for you

The discipline in all of this is unglamorous and mostly sequential: know what your collection contains before you specify how to photograph it, know what a reader will search for before you decide what text to generate, and build a QA gate that is genuinely capable of rejecting your own work. The technology available for reading difficult pages has improved substantially and will keep improving. Judgment about which series to read, in what order, to what standard, and how to prove it — that remains the archivist's.

Frequently Asked Questions

What are the stages of a mass digitization project?

A mass digitization project runs in a fixed sequence: collection survey, rights clearance, prioritization, capture, text generation, description and encoding, quality assurance, publication, and preservation. Each stage constrains the next, which is why the survey matters so much — the search outcome you intend dictates capture specification, the text-generation route, and how you sample for QA. The two stages institutions most often under-plan are generating the text layer and proving it worked. A page image is not searchable text; the text layer comes from a separate process, whether OCR, handwritten text recognition, human transcription, or a combination.

How much does capture cost as a share of a digitization budget?

The Northeast Document Conservation Center reports, cautiously, that high-resolution digital capture consumes perhaps one-third of the cost of a digitization project. That figure is experience-based rather than a current per-page price, but it works as a sanity check: if capture is your whole budget, you have not budgeted for the project. The remaining two-thirds covers descriptive metadata, text generation, quality assurance, publication, and storage. Planning that treats the scanner as the finish line tends to produce collections of images that nobody can search.

Can OCR read handwritten historical documents?

No — OCR does not read manuscript hands reliably, and planning as though it does is the most expensive misconception in digitization work. Handwriting needs handwritten text recognition (HTR) or human transcription. Historical print is also its own problem: hand-press printing produced fluctuating baselines, mixed founts, and uneven ink, so conventional OCR mishandles the long s, splits or drops ligatures, and silently normalizes archaic spelling. Route material by what it actually is — modern clean print to general OCR, historical print to tools customized for it, handwriting to HTR or people.

Do I need to train a custom HTR model for my collection?

Not necessarily, and the decision is worth making deliberately rather than by default. Training a custom model in Transkribus, eScriptorium, or OCR4all can produce strong results on a consistent series, but it front-loads ground-truth labour — manually finding and transcribing text lines is, as the peer-reviewed OCR4all paper puts it, a highly non-trivial task. Budget it as staff time against a real page count. A zero-shot specialist model removes the ground-truth phase from the critical path entirely, which suits collections with mixed hands or too little volume per series to justify training. Test a ready model on your own pages first.

How accurate is machine transcription, and what error rate is good enough?

There is no validated universal threshold for "good enough for search" in current practice, and anyone quoting you one is guessing. Set your own, tied to your intended use. Measure character error rate and word error rate against ground truth you create, on a random sample of representative pages, per series — and sample hard material separately, because an aggregate figure averaged across clean prose and ledger grids tells you nothing actionable about either. Weight the sample toward tokens that carry meaning: names, dates, places, sums. A 4% CER spread across function words is a different collection from 4% concentrated in surnames.

Share this article

© 2026 Leo Technologies Limited. All rights reserved