Digitization Staffing: Mapping the Roles, Functions, and FTE Estimates a Manuscript Project Actually Needs

Mapping digitization staffing for manuscript projects by enumerating required functions, assigning owners, and building FTE estimates from timed pilots on your own material rather than universal ratios.

Leo Team

August 18, 2026

Digitization Staffing: Mapping the Roles, Functions, and FTE Estimates a Manuscript Project Actually Needs
Contents

This is a working guide to digitization staffing for archives, libraries, and museums: how to enumerate the functions a manuscript project requires, assign an owner to each, and build FTE estimates from measurement on your own material rather than from someone else's ratio. It matters because the most expensive staffing mistakes are not made in the budget total — they are made in the functions nobody costed.

Digitization staffing is a function-mapping problem before it is a headcount problem. The defensible method is to enumerate every function the project requires — selection, condition and rights screening, preparation, capture, image processing, image QC, descriptive metadata, text generation and content verification, delivery, preservation, and project management — then decide, for your collection and your quality target, which of those functions a person owns, which are shared, and which are contracted out. There is no reliable universal pages-per-FTE ratio, and no authoritative source publishes one: the U.S. National Archives' digitization strategy states that digitizing for public access crosses multiple business units and requires a separate human-resource plan, without specifying numeric FTEs. Treat an FTE as a unit of staff capacity over a stated period and a stated task scope — never as a fixed page throughput.

That framing matters because the most common staffing failure is not underfunding. It is funding the wrong stage: budgeting a scanning operation and discovering, eighteen months in, that the images exist and nothing is findable.

Why "how many FTEs do I need?" has no clean answer

Every experienced project manager has been asked for a ratio. The honest response is that the literature defines functions and quality requirements, not staffing formulas. The FADGI still-image technical guidelines specify aim points and tolerances across measures such as tone response, white-balance error, illuminance uniformity, spatial-frequency response, noise, and sampling frequency — ranked on a four-star performance scale, with tighter tolerances meaning higher performance and higher cost. That tells you how good the images must be. It does not tell you how many people it takes to make them. The Metamorfoze preservation imaging guidelines are narrower still: they relate solely to the image quality of first-generation camera or scanner files. A capture and QC requirement, not a labour ratio.

The Digital Library Federation's Digitization Cost Calculator asks you to enter a number of scans and returns time and cost estimates drawn from contributed institutional data. It is useful for sanity-checking a capture budget. It is not evidence of a universal benchmark, and it does not model the text layer at all.

Nor is there strong comparative evidence about which stage is normally the largest cost for manuscript projects. FADGI and CLIR both establish that metadata and quality control require ongoing skilled human work — FADGI notes that metadata quality assessment will likely require skilled human evaluation rather than machine evaluation — but neither justifies a universal percentage. If a plan tells you that metadata is 40% of any digitization budget, ask where that number came from.

So the workable approach is to write the function list first, then attach capacity to it.

The function list

Adapted from the NARA enumeration and the FADGI concerns, a manuscript digitization project contains the following distinct functions. Some collapse into one person in a small institution. None of them disappear.

Selection and planning

Deciding what gets digitized, in what order, to what standard, and why. This is a curatorial and managerial judgment, not a technical one, and it is where the whole downstream cost is set. It belongs with someone who knows the collection and can defend the priority order to a funder or a board.

Condition, rights, and access screening

A conservator or collections-care specialist assesses condition, safe handling, preparation, and whether the chosen capture method risks the original. The Library of Congress's preservation scanning guidance describes minimal conservation treatment — "stabilization" — performed by trained conservators before items can be safely transported, handled, and digitized, and describes preservation staff training digitization technicians in careful handling using non-collection samples. Rights and access review runs in parallel: restricted material, third-party copyright, personal data.

Preparation and handling

Locating, pulling, unfolding, flattening, foliating, refiling. Unglamorous and consistently underestimated. In a series of bound volumes with fragile spines, preparation can exceed capture time per item.

Image capture

The imaging technician operates camera or scanner and maintains lighting, sampling, colour and tone, resolution, file, and handling discipline. FADGI is clear that production staff need a foundation in photography and imaging and real technical experience — this is a skilled role, not a keying role. Whether it is one FTE or three depends entirely on the capture standard you set and the format mix.

Image processing

Cropping, deriving, naming, packaging. Increasingly automated, but automation still requires a person to configure and audit it.

Image QC

Completeness, focus, colour and tone, resolution, distortion, file integrity, conformance to the target standard. This can sit with the capture team or be separated; the function must exist either way.

Descriptive metadata

The cataloguer or descriptive archivist creates or maps descriptive, technical, structural, and preservation metadata, works with controlled vocabularies, and situates items in catalogue context. If you are deciding how deep to go here, the layer model in our guide to archival metadata standards for transcribed documents is a better starting point than a single schema.

Text generation and content verification

The OCR/HTR operator, ground-truth producer, and verifier. Model selection and evaluation, running the recognition, comparing output against the image, correcting errors, applying language and paleographic judgment. This is a distinct skill from image QC and cannot be staffed from the same pool.

Delivery, preservation, IT

Repository ingest, storage, fixity, access platform. Usually shared with existing institutional infrastructure, which is why its true cost is so often invisible in project budgets.

Project management

Scope, budget, staffing and training, contractor management, space, dependencies, acceptance criteria. FADGI treats project management, staffing and training, image inspection, acceptance and rejection, metrology, and metadata quality management as separate concerns. A project of any size that treats management as a spare-time duty for a curator will pay for it in rework.

Where the work actually sits: four things staffing plans get wrong

QC is not a clerical final check

CLIR's framework for assessing preservation aspects of large-scale digitization is direct on this: quality control covers images, OCR output, and metadata, and although automated tools exist for file naming and integrity checks, some quality elements — missing pages, imaging distortions — can be detected only through visual inspection. Most quality assurance is still done manually.

The same source records how QC practice has shifted under volume: early library projects often ran 100% QC with visual comparison of digital and print pages, looking for wavy patterns, banding, Newton's rings. At institutions converting on the order of 10,000–40,000 books a month, exhaustive QC stopped being viable, and the field moved toward sampling strategies chosen against budget, infrastructure, staff qualifications, materials, and timeline. That is a decision your project has to make explicitly, with a documented sampling rate — not a decision made by default when the QC person runs out of hours.

And split the QC function in two when you write the job descriptions. Image QC needs a careful eye and a technical checklist. Content QC needs someone who can read the hand.

The verifier is a language-and-hand role, not a typing role

This is the staffing line most often mis-specified. Reading a seventeenth-century Dutch notarial register is not a generic clerical competence. The verifier needs the language, the historical orthography, the abbreviation repertoire, and the local record tradition — German Kurrent, English secretary hand, French notarial cursive, each with its own hazards. A 2025 assessment of advanced HTR engines makes the point implicitly by reporting how performance shifts with language, script, orthography, document condition, and dataset; one of its corpora consists of seventeenth-century printed Dutch and French texts, which the authors note present particular challenges from historical orthography and multilingual content.

A related confusion worth heading off in a staffing memo: Latin script is a writing system, not a language. Saying your collection is "Latin-script" says nothing about whether your verifier needs Latin. Most vernacular archival series — parish registers, wills, deeds, minutes — are written in the local language using the Latin alphabet, and it is that vernacular competence you are hiring for. Our guide to which languages and scripts HTR can read works through why script, language, and hand are three separate variables.

Budget the verifier by hand difficulty and record series, not by page count across the whole project. A clean nineteenth-century register and a crabbed early-modern letter book are not the same job.

HTR moves labour; it does not delete it

If a plan proposes cutting transcription staff because the project has adopted machine transcription, the arithmetic is probably wrong. Recognition removes keystrokes. It does not remove verification, editorial judgment, structural correction, or metadata work — and in tool-training workflows it adds ground-truth production.

How much it adds depends on the tool class. Trained-model platforms make the ground-truth step a staffing line in its own right: eScriptorium's training documentation describes an iterative loop — correct two or three pages of automatically generated transcriptions, fine-tune, evaluate, correct more, repeat — while training from scratch requires substantial and diverse data. Transkribus similarly requires ground-truth pages before a text-recognition model can be trained. OCR4all's workflow documents preprocessing, layout analysis, ground-truth production, training, recognition, and post-correction as distinct stages, and is a print-OCR pipeline rather than a handwriting one. None of these are objections to the tools; they are staffing facts. Someone has to do the correcting, and that someone needs the same paleographic competence as your verifier. If you are weighing this, our piece on whether you need to train your own model sets out the conditions under which training earns its cost.

The post-correction curve is also non-linear in a way worth planning for. A user-perspective study of HTR on the Codex Runicus recorded manual transcription at 14–21 minutes per page including validation, and reported validation totals across four pages of 30 minutes for one machine-assisted method and 29 for another, against 64 minutes for fully manual transcription — but the authors also note that where character error rate rose, validation time rose with it, approaching the cost of transcribing manually (Codex Runicus study). That is a small, specialist rare-script case, not a vernacular archival benchmark. The transferable lesson is structural: verification effort scales with error rate, so recognition quality is a staffing variable, not merely a quality variable. A model that reads your hands well converts a transcription post into a checking post. A model that reads them badly converts it back again.

Crowdsourcing substitutes a coordination model for a payroll

Volunteer transcription is a legitimate and often excellent choice, but it is not free labour, and it does not remove staffing from the plan. Transcribe Bentham's volunteers produced 1,009 transcripts — an estimated 250,000 to 750,000 words plus mark-up — during a six-month testing period, with seven "super transcribers," 0.6% of registered users, accounting for 709 of them (Causer et al., DHQ). That concentration is the operational reality of volunteer programmes: output depends on a small, cultivated core. The same team's later cost-effectiveness analysis examined 4,364 checked and approved transcripts submitted between October 2012 and June 2014, reported 94% of transcribed or partially transcribed manuscripts approved at the time of writing, and estimated potential cost avoidance of roughly £500,000 after accounting for £589,000 already invested — a project-specific scenario, not a general return.

The Library of Congress's By the People programme runs a consensus model in which two or more volunteers must agree before a transcription is complete, with subject specialists spot-checking completed campaigns before publication. Consensus review, task design, interface maintenance, community management, promotion, and specialist checking are all staffed functions. Budget a community coordinator and a specialist reviewer, or the queue stalls. Our companion piece on what crowdsourced transcription volunteers do best works through the series-by-series decision.

Building the estimate

Since no ratio exists, build capacity from measurement on your own material.

  1. Write the function list above against your project, and name an owner for each line — including the ones you intend to share, contract, or defer. An unowned function is a schedule risk.
  2. Set the quality targets before the staffing. Which FADGI star level, for which measures? Item-level or collection-level description? Verbatim transcription or searchable full text at a stated error tolerance? Each target sets labour.
  3. Run a timed pilot on a representative sample — twenty to fifty items drawn from each distinct record series and hand. Time preparation, capture, image QC, description, recognition, and verification separately. This is the only staffing evidence that will hold for your collection.
  4. Convert to FTE with a stated denominator. "0.5 FTE verifier for eighteen months, at the measured rate of X pages per hour for series Y, at a 5% sampling QC rate" is defensible. "2 FTE for digitization" is not.
  5. Re-measure after the first batch. Rates move as staff learn the hands and the tooling settles.

For the sequencing that sits underneath this — how preparation, capture, QC, metadata, and text recognition constrain each other — the stage-by-stage archival digitization workflow guide covers the pipeline, and the broader planning context sits in our overview of archives, digitization and metadata workflows. If you are costing this rather than staffing it, what a records digitization project actually costs breaks out the line items.

The text layer: where the tool choice changes the staffing line

Because verification effort tracks recognition quality, the transcription engine you pick is a staffing decision. Two properties matter more than feature lists.

The first is whether the tool requires you to build a model before it reads anything. If it does, ground-truth production becomes a standing post staffed by someone with paleographic competence — the scarcest resource in most institutions. A zero-shot model that reads your hands adequately out of the box removes that post, or defers it until you have evidence you need it.

The second is how the tool fails. Errors that look like errors are cheap to catch; a verifier scanning output against the image beside it spots a garbled word in a second. Errors that read fluently and plausibly are expensive, because catching them requires reading the manuscript properly — which is the labour you were trying to reduce. That is the specific reason general chatbots are a poor foundation for a production text layer, however capable they seem on a single page: they smooth archaic orthography into modern prose and fill gaps with plausible readings. Our piece on fluent but wrong LLM transcription errors documents the failure mode.

This is the stage Leo is built for. ATR-1, Leo's transcription model, is zero-shot on Latin-script material of roughly the past 500 years — English, French, German, Dutch, Spanish, Italian and other languages written in that alphabet, handwritten or printed — so there is no per-collection training step and no ground-truth post to staff before the first page is read. It transcribes what is on the page: strikethroughs, insertions, marginalia, archaic spelling, and expansions are preserved rather than silently normalized, which is what a verifier needs in order to check output against the image rather than against their expectations. On a randomized 97-image sample of early-modern English manuscripts from the Folger Shakespeare Library at ATR-1's release, Leo scored approximately 5% character error rate against roughly 13% for Transkribus's Text Titan I and 23–25% for Claude Opus and Gemini 2.5 Pro (full benchmark data) — a single-corpus test, and your own timed pilot on your own series is still the number that should drive your staffing line. Around the model sits the workflow the verifier actually lives in: folders, per-document archival metadata, image beside transcription, fuzzy search across everything, and TEI, Word, PDF or HTML export. Two honest limits for planning purposes: pages that mix dominant printed structure with dense handwriting — pre-printed ledger and deed-book forms — are a known weak spot, and non-Latin scripts are out of scope entirely. Academic and archival projects can apply for up to 100,000 free credits through the Leo Transcription Grant, on condition that the resulting transcriptions and images are published openly — which means applying with material you hold or can clear the rights to publish.

Staff the functions you can defend

The value of a staffing plan is not the total headcount. It is that every function has a named owner, every quality target has a measurement behind it, and every rate came from your own material rather than from someone else's project. That plan will survive a funder's questions, a change of scope, and the discovery in month four that the bound volumes take three times longer to prepare than the loose files.

It also gives you something to revise honestly. Measure the first batch, adjust, and record why. Digitization programmes that last are the ones that treat their own throughput data as evidence — the same discipline the collections themselves deserve.

Frequently Asked Questions

How do you plan digitization staffing for a manuscript project?

Start by enumerating functions, not headcount. List every function the project requires — selection and planning, condition and rights screening, preparation and handling, image capture, image processing, image QC, descriptive metadata, text generation and content verification, delivery and preservation, and project management — then name an owner for each line, including the ones you intend to share, contract, or defer. Set quality targets next, run a timed pilot on twenty to fifty items from each distinct record series and hand, time each stage separately, and convert to FTE with a stated denominator. Re-measure after the first batch.

Is there a standard pages-per-FTE ratio for digitization?

No, and no authoritative source publishes one. The U.S. National Archives' digitization strategy states that digitizing for public access crosses multiple business units and requires a separate human-resource plan, without specifying numeric FTEs. FADGI defines image quality aim points and tolerances on a four-star scale; Metamorfoze addresses first-generation file quality only. Both tell you how good the output must be, not how many people it takes. Treat an FTE as staff capacity over a stated period and task scope, never as a fixed page throughput, and build rates from measurement on your own material.

What skills does a transcription verifier need?

Language and paleographic competence, not typing speed. Reading a seventeenth-century Dutch notarial register is not a generic clerical task: the verifier needs the language, the historical orthography, the abbreviation repertoire, and the local record tradition — German Kurrent, English secretary hand, French notarial cursive each carry their own hazards. Note also that Latin script is a writing system, not a language; most vernacular series such as parish registers, wills, deeds and minutes use the Latin alphabet in the local language, and it is that vernacular reading ability you are hiring. Budget by hand difficulty and record series, not total page count.

Does HTR reduce digitization staffing costs?

It moves labour rather than deleting it. Recognition removes keystrokes but not verification, editorial judgment, structural correction, or metadata work, and on trained-model platforms it adds ground-truth production as a staffing line of its own — eScriptorium and Transkribus both require corrected pages before a model can be trained. Verification effort also scales with error rate: a user study of HTR on the Codex Runicus found that as character error rate rose, validation time approached the cost of transcribing manually. A model that reads your hands well converts a transcription post into a checking post; a poor one converts it back.

Is crowdsourced transcription a cheaper alternative to paid staff?

It substitutes a coordination model for a payroll rather than removing staffing. Transcribe Bentham's volunteers produced 1,009 transcripts during a six-month testing period, with seven "super transcribers" — 0.6% of registered users — accounting for 709 of them, which shows how output depends on a small, cultivated core. A later cost-effectiveness analysis of 4,364 approved transcripts estimated roughly £500,000 in potential cost avoidance after £589,000 already invested — project-specific, not a general return. Consensus review, task design, interface maintenance, community management and specialist spot-checking are all staffed functions. Budget a coordinator and a specialist reviewer.

Share this article

© 2026 Leo Technologies Limited. All rights reserved