How to Clear a Digitization Backlog: Triage, Batching, and Where Crowdsourcing Runs Out

Clearing digitization backlogs by diagnosing which queue each series is stuck in, batching by hand and layout, setting fidelity by use, and placing machines and crowdsourcing only where each fits.

Leo Team

August 18, 2026

How to Clear a Digitization Backlog: Triage, Batching, and Where Crowdsourcing Runs Out
Contents

A digitization backlog is rarely one queue, and it is almost never solved by more scanning or more volunteers. This is a practical method for diagnosing where each series is actually stalled, batching material by hand and layout, setting a fidelity target you can defend, and deciding when a machine first pass earns its place in the pipeline.

A digitization backlog is a chain of queues — selection and rights, capture, storage, description, text recognition or transcription, indexing, and publication — and material stalls in different places for different reasons. Clearing it means diagnosing which queue is blocked for each series, then setting a fidelity target proportionate to how that series will be used: minimal description and image access for some, machine-generated searchable text for others, verified transcription only where the use case demands it. The common failure is treating the whole backlog as a scanning problem, or as a problem more volunteers would solve.

Locate the blockage before you buy a solution

The word "backlog" flattens a set of very different problems. A shelf of unopened accessions and a hard drive of 40,000 unlinked TIFFs are both backlogs, and almost nothing you do about one helps the other.

The most useful sector framing is OCLC's: a hidden collection is special-collections or archival material that is undescribed or under-described, and therefore undiscoverable. Note what that definition does not mention. Scanning. A collection can be fully imaged, stored, and backed up, and remain hidden, because nothing about it is findable. OCLC's 2010 Taking Our Pulse survey — 275 libraries surveyed, 169 responding — reported that half of archival collections had no online presence, and that while many backlogs had decreased, almost as many continued to grow. The figure is old and survey-specific, and it does not establish a universal ratio of imaging debt to description debt. What it does support is the more careful diagnosis: capture produces surrogates faster than description, transcription, rights review, and indexing can make them usable.

So the first act of backlog clearance is an audit, series by series, that answers four questions:

  • Is it imaged? If not, is that because of cost, fragility, equipment, or because nobody has decided it is worth imaging?
  • Is it described? At what level — collection, series, item — and is that level sufficient for the way people ask for it?
  • Is it readable? Are there transcripts, machine or human, and is anything full-text searchable?
  • Is it publishable? Rights, donor restrictions, privacy, and sensitivity work is its own queue, and it blocks material that is otherwise finished.

Most institutions find that these four are blocked in different proportions across their holdings. That is the point. A single procurement — a scanner, a vendor contract, a transcription platform — will address one column of the audit and leave the others exactly as they were.

Triage: rank by demand, coherence, and required fidelity

Greene and Meissner's More Product, Less Process argument remains the right instinct at the top of the pipeline: proportionate processing exposes more holdings sooner than exhaustive item-level work. Their processing-rate study found manuscript collections under a cubic foot took 5.5 days per foot, larger manuscript collections 3 days per foot, and archival series of any size 2 days per cubic foot. The lesson is not "describe less." It is that description effort should be sized to the material and its likely use — and the same logic extends past description into the text layer.

Rank each candidate series on five axes.

Demand

Reference requests, reproduction orders, citations in published work, and repeated enquiries from local family historians are your cheapest signal. Series nobody has asked for in a decade can wait behind series that generate weekly correspondence.

Rights and sensitivity

A series that cannot be published is a poor candidate for a project justified by public access, however easy it is to image. Clear this early; it is the queue most often discovered late.

Legibility

Faded iron-gall ink, heavy show-through, foxing, and water damage change both capture requirements and recognition outcomes. This is a distinct diagnosis from handwriting difficulty, and worth separating — some of it is a capture problem rather than a recognition problem, and a few are conservation problems that no imaging decision will fix.

Homogeneity

This is the axis most often skipped and the one that most determines whether machine transcription will work. A register kept by three clerks over forty years in one house style is a very different proposition from a correspondence file with two hundred correspondents. Homogeneous series are the best first-pass candidates for automated text; miscellaneous series are not.

Required fidelity

Ask what the text is for. A finding-aid enhancement that lets a researcher locate a name needs different accuracy than a text that will be quoted, cited as legal evidence, or published as an edition. Transcription, translation, and a scholarly edition are three separate deliverables, and conflating them inflates every estimate you make. If your institution has not drawn the line between finding-aid metadata and full-text transcription, that distinction is worth settling before the pipeline is designed.

The output of triage is not a yes/no digitization list. It is a ladder: minimal description plus image access for the bulk of holdings; machine-assisted searchable text where the material is coherent and the demand is real; human verification concentrated on the fields and passages that carry weight; full editorial transcription for the small subset that genuinely warrants it.

Batching: group by hand and layout, not by shelf order

Backlogs are usually organized by provenance and stored by accession. Neither is the right unit for processing work.

Batch instead by the characteristics that determine effort. A batch should share a script family, a broad period, a layout type, and a fidelity target. Practically, that means separating:

  • Bound registers with a stable layout — parish registers, minute books, deed volumes — where every page repeats a known structure.
  • Loose correspondence and papers — variable hands, variable page geometry, marginalia.
  • Tabular material — musters, censuses, ledgers, station registers — where structure carries as much meaning as the characters.
  • Printed matter — pamphlets, broadsides, newspapers, early-modern books — which has its own failure modes, and which is often mixed into manuscript collections without being flagged.
  • Mixed print-and-manuscript forms — pre-printed ledger and certificate forms completed by hand — which are the hardest category for any automated system and should be scheduled deliberately rather than swept in with the rest.

Batching this way lets you sample properly. Before committing a batch, take a small representative set of pages — a genuine random sample, not the legible ones — and produce careful reference transcriptions against which output can be measured. This is your ground truth: the trusted reference transcription used to evaluate, and in trained-model workflows to train. Then measure character error rate on held-out pages that were not used to tune anything. CER is the proportion of character-level edit errors against that reference; word error rate applies the same idea at word level. Both are only as meaningful as the sample they were computed on, and how to run this test on your own material in an afternoon is a skill worth having in-house rather than outsourcing to a vendor's demo.

Do not carry one batch's result to another. In the 2025 multi-engine benchmark by Romein and colleagues, reported figures ranged from 1.50% base CER on seventeenth-century Dutch/French Roman-type print to 9.29% on German shorthand, with Dutch States General resolutions at 1.78% for one engine. Those numbers are not interchangeable, and the print result is not evidence about handwriting. Performance is collection-, hand-, layout-, period-, and language-dependent, and the paper is candid about its own dataset limitations, including one ground-truth set initially generated by an engine and then corrected by hand.

Where crowdsourcing runs out

Crowdsourced transcription is not a low-quality option. Run properly — with image and handwriting guidance, repeated independent passes or consensus rules, editable records, and staff moderation — it produces work that stands up. Transcribe Bentham's quality and cost analysis reports that 94% of transcribed or partly transcribed manuscripts had been approved by project staff. It also builds a public constituency for the collection, which no software does.

The constraint is throughput, and its shape is worth understanding precisely. In the project's six-month test from September 2010 to March 2011, 1,207 people registered and 259 — 21% — actually transcribed, producing 1,009 manuscripts and over half a million words with mark-up. Seven "super transcribers," 0.6% of registered users and 3% of those who transcribed, worked on 709 of those manuscripts: 70% of the total. The same analysis estimated that two full-time research associates, had they spent those six months transcribing, could reasonably have produced around 2,400 transcripts — more than twice the volunteer rate. That counterfactual excludes the outreach value, and it excludes the fact that those associates were, in reality, doing the moderation that made the volunteer work usable.

Read that as a distribution, not a verdict. Volunteer capacity concentrates in a small core, recruitment and retention are ongoing costs, and community management is staff labour that does not disappear when you add volunteers. A programme that depends on seven people is a programme with a succession risk. And these are one project's statistics, not a universal law; comparable person-hour and retention data across Zooniverse, FromThePage, By the People, and the Smithsonian Transcription Center are genuinely scarce. What the evidence supports is a placement rule rather than a replacement rule: use volunteers where the engagement and quality-control economics fit, and do not budget them as an elastic workforce for a million-page series.

Where crowdsourcing runs out is at scale on homogeneous, unglamorous material. No volunteer community sustains enthusiasm through 300,000 pages of nineteenth-century notarial deeds. That is exactly the material a machine first pass handles best — and exactly the material where volunteers, redeployed to verification of names, dates, and figures, add the most value per hour.

The economics of the first pass

The decision to run automated transcription rests on one question: is checking a machine draft faster than reading the page from scratch?

The Codex Runicus evaluation puts numbers on this. Manual transcription there took 14–21 minutes per page; validating recognition output took 5–7 minutes per page against 14–15 minutes manually, with one page recorded at 21 minutes manual versus 5 or 7 minutes to validate. The authors are explicit about the limit: as error rates rise, validation time approaches manual time, and above a certain error burden correction time can exceed transcribing from scratch. That is a rare script and a small study, not a benchmark for vernacular Latin-alphabet backlogs, but the shape of the curve transfers. A good first pass saves most of the reading time. A bad first pass costs more than no first pass at all.

Which is why the sampling step is not optional overhead. It is the measurement that tells you which side of that curve a batch sits on before you commit thousands of pages to it.

Scale is achievable. The National Archives of the Netherlands processed 3 million pages — chiefly seventeenth- and eighteenth-century Dutch East India Company records and nineteenth-century notarial deeds — in a 2020–21 project, reporting 7% CER after building 6,000 pages of training data against an initial 20% target. Two things about that account deserve attention. It is a vendor case study, not an independent evaluation: no WER, no recall or precision, no end-to-end cost. And the team's own remark is the most useful line in it — transcription was the easy part; publishing everything online was substantially harder, archivally and technically. Budget accordingly.

Note also what 6,000 pages of training data represents. In trained-model workflows — Transkribus, eScriptorium with Kraken, OCR4all — you must find or build a model before you get output, and building one requires correctly segmented examples and trusted ground truth. For a large homogeneous series that investment amortizes. For twelve small series with different hands, it does not, and whether you need to train your own model at all is worth testing before it becomes a project assumption.

Where Leo fits in this stage

Leo is built for this position in the pipeline: after capture, before publication. Its transcription model, ATR-1, reads Latin-script material — whatever language is written in that alphabet, English, French, German, Dutch, Spanish, Italian, Latin among them — with no per-collection model training and no segmentation step, which is what makes it practical for an archive with a dozen heterogeneous series rather than one large uniform one. It reads printed matter as well as manuscript hands, including early-modern typography that conventional OCR mishandles: the long s, ligatures, typographic abbreviation, show-through read as character evidence. Non-Latin scripts — Greek, Cyrillic, Hebrew, Arabic, Indic, East Asian — are out of scope.

On accuracy, the comparison we publish is a randomized 97-image sample of early-modern English manuscripts from the Folger Shakespeare Library, measured at ATR-1's release: Leo at roughly 5% CER, against approximately 13% for Transkribus's Text Titan I, 23.3% for Claude Opus, 24.8% for Gemini 2.5 Pro, and 56.7% for GPT-4.1 — 61% fewer errors than the next-best model, with the full results published here. That is one corpus, one hand tradition, one date. Sample your own material.

Two properties matter specifically for backlog work. First, source integrity: ATR-1 transcribes what is on the page — strikethroughs, insertions, marginalia, archaic orthography — rather than normalizing it into modern prose, which keeps the machine text defensible as a record and keeps verification a matter of checking characters rather than hunting for fluent inventions. Second, the workflow around it: uploads of a thousand-plus images at once, per-document metadata, folders, fuzzy search across everything transcribed, TEI, Word, PDF and HTML export, and public read-only share links for a document or a whole folder. Transcription stays separate from interpretation — translation, summarization, entity extraction and the rest run as Transformations that write to new tabs, leaving the base transcription untouched.

The limits are worth stating plainly. Pages mixing dominant printed structure with dense handwriting — pre-printed ledger and deed-book forms — are the known weak spot, and complex tabular layouts vary. Schedule those batches for heavier human review, or defer them. Leo begins at upload; it does not scan, and it does not do IIIF, OAIS packaging, or EAD/METS/PREMIS, which your repository system handles. Academics and archives can apply for the Leo Transcription Grant — up to 100,000 credits against a commitment to publish the transcriptions and images openly, which means applying with material you hold or can clear the rights to publish.

Publish before it is finished

The last thing that keeps backlogs closed is the instinct to wait for completeness. Machine text clearly labelled as machine text, sitting beside the image, with a correction route, serves researchers today. Held back until it is verified, it serves nobody.

Label the provenance of every text layer — model output, volunteer transcription, staff-verified, published edition — and record corrections against it. Route your scarce human attention to the tokens that carry consequence: personal and place names, dates, sums, boundary calls. Those are where an error propagates into a family tree, a title chain, or a footnote, and where targeted verification buys more than a uniform read-through ever will. The rest of the page can improve incrementally.

None of this is a formula. It is the ordinary archival judgment you already exercise — what this collection is, who asks for it, what it must be good enough for — applied one stage further down the pipeline than description has traditionally reached. The backlog does not clear in a single project. It clears series by series, each with a fidelity target you chose deliberately and can defend, and each published the moment it is useful rather than the moment it is finished. That, more than any tool, is what turns a shelf of hidden holdings into a collection people can actually read. For the wider view of how these stages constrain one another, the archives, digitization and metadata workflows material covers the pipeline end to end.

Frequently Asked Questions

How do you clear a digitization backlog?

Start by diagnosing where each series is actually stalled, because a digitization backlog is a chain of queues — selection and rights, capture, storage, description, text recognition, indexing, publication — not one blocked pipe. Audit series by series: is it imaged, described, readable, publishable? Then triage on demand, rights and sensitivity, legibility, homogeneity, and required fidelity, batch by hand and layout rather than shelf order, sample before committing, and set a fidelity target proportionate to use. Publish each series the moment it is useful rather than the moment it is finished. Backlogs clear series by series, not in one project.

What is a hidden collection in an archive?

A hidden collection, in OCLC's framing, is special-collections or archival material that is undescribed or under-described and therefore undiscoverable. Notably, the definition says nothing about scanning. A collection can be fully imaged, stored and backed up and still be hidden, because nothing about it is findable. This is why treating a backlog as a scanning problem misdiagnoses it: capture produces surrogates faster than description, transcription, rights review and indexing can make them usable. OCLC's 2010 Taking Our Pulse survey found half of archival collections had no online presence, with many backlogs still growing.

Can volunteers clear a large transcription backlog?

Not at scale on homogeneous material. Crowdsourced transcription, run with proper guidance, consensus rules and staff moderation, produces work that stands up — Transcribe Bentham reported 94% of transcribed or partly transcribed manuscripts approved by staff. But its six-month test saw only 21% of registrants transcribe, with seven "super transcribers" producing 70% of the manuscripts, and two full-time research associates were estimated to have produced more than twice the volunteer output. Volunteer capacity concentrates in a small core, and community management is staff labour. Use volunteers where engagement and quality-control economics fit, not as an elastic workforce.

When is a machine first pass worth running on archival material?

When checking the machine draft is faster than reading the page from scratch. The Codex Runicus evaluation recorded manual transcription at 14–21 minutes per page against 5–7 minutes to validate recognition output — but its authors are explicit that as error rates rise, validation time approaches manual time, and beyond a point correction costs more than transcribing fresh. That is why sampling matters: take a genuine random sample, produce careful reference transcriptions, measure character error rate on held-out pages, and find out which side of that curve a batch sits on before committing thousands of pages.

How should archives batch material before transcription?

Batch by the characteristics that determine effort, not by provenance or accession order. A batch should share a script family, a broad period, a layout type and a fidelity target. In practice that means separating bound registers with stable layouts, loose correspondence with variable hands and page geometry, tabular material where structure carries meaning, printed matter with its own failure modes, and mixed print-and-manuscript forms such as pre-printed ledgers and certificates — the hardest category for any automated system, and best scheduled deliberately for heavier human review or deferred. Batching this way also lets you sample each group properly.

Share this article

© 2026 Leo Technologies Limited. All rights reserved