How to Cite AI Transcription in Scholarly Work: Method Notes, Disclosure, and What Reviewers Ask

How historians should cite AI transcription in published scholarship: the manuscript takes the citation while HTR is disclosed in method notes covering platform, verification, and editorial decisions.

Leo Team

September 17, 2026

Contents

This is a practical guide to how to cite AI transcription in published historical scholarship: what belongs in the citation, what belongs in a method note, and what reviewers ask when handwritten text recognition has touched your sources. The short version is that the manuscript takes the citation and the recognition system is disclosed as a method — but the details of that disclosure are what make the work defensible.

Cite the manuscript, not the machine. When handwritten text recognition (HTR) produces a transcription you quote in published work, the archival document remains the evidentiary source and takes the citation; the recognition system is a method, disclosed in a methods paragraph, an editorial note on transcription, or a footnote at first use. A defensible disclosure names the platform and model or version, the date it was run, the images and pages it was applied to, whether any model training was involved, how much human verification was done against the original image, and any downstream correction, normalization, or translation. The generative-AI citation formats issued by Chicago, MLA, and APA were written for text a model composed, and they transfer badly to text a model read off a page you can still check.

That last point is where most confusion starts, so it is worth working through properly.

Why the generative-AI formats do not fit

The major style guides responded quickly to ChatGPT and have kept their guidance current. Chicago's Q&A on citing content generated by AI offers a note in the form "Text generated by ChatGPT, OpenAI, March 7, 2023," treating the tool as a stand-in author and the developer as publisher — and advises against a bibliography entry unless a publicly available link exists, on the grounds that a private conversation is closer to personal communication. The MLA Style Center declines to treat the tool as an author at all, routing generated material through its core-element template: describe what was generated, name the tool as container, give the version, the company, the generation date, and the general URL (an updated version of that guidance was announced for August 2025). APA treats the model as a software work: OpenAI. (2023). ChatGPT (Mar 14 version) [Large language model], with the URL.

Three style guides, three answers — and all three are answering the same underlying question: how does a reader get at the thing you are quoting, when the thing you are quoting exists only because you prompted for it? Generated text has no prior existence. It is not retrievable. It has to be identified, dated, and evaluated as model output.

HTR is a different operation with a different evidentiary structure. The text existed before the model ran. It exists on a leaf in a box in an archive, and it will still be there when a reader goes to check. The recognition pass produced a reading of an existing artefact; it did not produce the artefact. So the citation obligations of the manuscript itself — repository, collection, series, box, folio, and the other retrieval coordinates a defensible archival footnote must carry — do not change at all. What changes is that you now owe your reader an account of how the reading was produced.

No major style guide sets out a dedicated rule for HTR or transcription. There is no template to copy. What exists instead are two well-established traditions you can extend: software citation and documentary editing.

Source versus method: the distinction that decides everything

Treat the recognition system the way you would treat any other software dependency in a research pipeline. FORCE11's Software Citation Principles make the reproducibility argument plainly: identifying the software you used enables better peer review. That is the whole case, and it applies to an HTR engine exactly as it applies to a statistical package.

Version identifiers matter for the same reason they matter elsewhere. The Tesseract repository documents major version 5 as the current stable line, with releases continuing beneath it; eScriptorium is documented as a workspace for managing the steps of a transcription campaign, with segmentation and recognition supplied by Kraken components. A reader who knows only "we used HTR" knows nothing reproducible. A reader who knows the platform, the engine, the model identifier, and the date knows what you actually ran.

Hosted models complicate this, and it is better to say so in your note than to pretend otherwise: a cloud model can be updated without a version bump visible to you. Record the date you ran the pass. The date is the version stamp you can always defend.

What a defensible method note contains

Nine elements. Not all will apply to every project; state the ones that do.

  1. Platform and engine. The application you worked in and the recognition model that produced the text, distinguished if they differ.
  2. Model identifier and version, or the run date. Whichever you can pin down; ideally both.
  3. Image conditions. What was recognized — your own reading-room photographs, repository surrogates, microfilm scans — and at what resolution. Capture quality bounds recognition quality, which is why how the images were made belongs in the note.
  4. Scope. Which documents, which pages, how many images. If some series were transcribed by hand and others by machine, say which.
  5. Training and ground truth, if any. If you trained or fine-tuned a model, describe the ground-truth set and its size. If you used a ready model with no training step, say that — it is a shorter and cleaner thing to disclose, and reviewers read it as such.
  6. Human verification. The heart of the note. Was every page checked against the image, or a sample? By whom? What was prioritized — names, dates, sums, legal formulae?
  7. Evaluation figures, if you measured any. With conditions attached (below).
  8. Downstream processing. Correction, modernization, expansion of abbreviations, translation. Each is a separate editorial act, not part of the transcription.
  9. The uncertainty key. What your brackets mean.

A worked note, in the register a monograph or article would actually use:

Transcriptions from TNA SP 16/234–236 were produced from colour photographs taken by the author in June 2025. A first-pass machine transcription was generated with [platform, model identifier], run 12 September 2025; no model was trained or fine-tuned for this project. Every page was subsequently checked against the image by the author. Original spelling, capitalization, and punctuation are retained; abbreviations are expanded in square brackets; readings that remain uncertain are marked [?] and wholly illegible passages [illegible]. Translations from French are the author's, made from the verified source-language transcription and not from machine output.

Roughly one hundred words. It answers every question a careful reader would raise, and it commits you to nothing you cannot defend.

Where the disclosure goes

This is genuinely unsettled. Integrity bodies speak in terms of generative or AI-assisted technologies rather than recognition specifically. COPE's position states that AI tools cannot be listed as authors, and that authors using AI tools in writing, in producing images, or in the collection and analysis of data must disclose in Materials and Methods or a similar section how the tool was used and which tool it was. ICMJE requires disclosure at submission and description in the manuscript, and likewise bars AI authorship. Cambridge's AI research-ethics policy requires AI use to be declared and clearly explained, "just as we expect scholars to do with other software, tools and methodologies." Oxford University Press asks for tool and version, where and how it was used, why, and how the output was validated — while noting its guidance will keep changing.

None of these creates an HTR exemption for fully proofread work, and none names transcription explicitly either. A 2024 BMJ audit of publisher and journal instructions found that the AI guidelines it identified referred to generative models or generative ability rather than AI broadly — the policies were written with composition in mind, not recognition.

The practical answer: put the substance in one place and cross-refer. For an article, a methods or sources paragraph. For a monograph, a "Note on transcription" in the front matter, which is where documentary editors have always put it. Add a submission-stage declaration if the publisher asks for one. Do not scatter half the information into an acknowledgement and the other half into a footnote on page 214.

And do not omit it because you checked everything. Full verification is a fact you disclose, not a reason not to disclose. It is also the single strongest sentence in the note.

Marking uncertainty so the marks mean something

A method note that says "uncertain readings are bracketed" without a key has told a reader nothing. Editorial tradition already distinguishes the states cleanly. TEI provides `<unclear>` for a reading that cannot be transcribed with certainty, `<supplied>` for text supplied by the transcriber or editor — because of damage or an obvious scribal omission — `<gap>` for omitted material, `<sic>` for an apparently erroneous source reading, `<corr>` for a correction, and `<choice>` to group alternative representations such as source and regularized forms; `@resp` records responsibility and `@cert` carries certainty. Leiden conventions draw the same distinctions in plain text, separating supplied, restored, omitted, and corrected readings by bracket type.

Two things follow. First, state your editorial policy: a diplomatic transcription represents the features of the source, a normalized one regularizes selected features, and the edition must say which it uses and where it departs. Second, decide whether machine-derived uncertainty enters the encoding at all. Mapping an HTR confidence value or a verification status onto `<unclear cert="low">` is a reasoned application of existing infrastructure, not a settled cross-field standard — so if you do it, define it. The broader conventions for what to leave exactly as written are unchanged by the presence of a machine in the workflow; the machine simply makes the policy statement more necessary, not less.

Reporting accuracy without overclaiming

Character error rate is character-level edit distance against ground truth, divided by the reference character count. It is an evaluation figure produced under specific conditions, not a property of a tool. It is conditional on the hand, the script, image quality, segmentation, the quality of the ground truth itself, and the sample — and the Romein et al. evaluation of advanced HTR engines is explicit about residual errors in ground truth and about the restriction of one engine's evaluation to Latin-script texts, which is a claim about the writing system rather than about the Latin language.

So: report a vendor's published benchmark as a benchmark, with its corpus named, and never as your project's accuracy. If you want a number that means something for your material, measure it on your own pages — a randomized sample, keyed by hand as reference, scored honestly — and state the sample size and composition alongside the figure. A short, conditioned local figure is worth more to a reviewer than any headline. It also helps to understand what these metrics do and do not measure before you put one in print.

There is one asymmetry worth naming, because it changes how much verification you should claim. Garbled output announces itself. Fluent output does not. A general chatbot asked to read a secretary hand will produce plausible, well-formed prose that silently smooths what is actually on the page — errors that survive proofreading precisely because they read well. If your first pass came from a general model, your disclosure needs to be correspondingly more detailed about verification, and your verification correspondingly more line-by-line.

Keeping a record you can write the note from

Most of the difficulty in writing these notes is retrospective: six months after the fact, nobody remembers which pages were machine-read, which were corrected, which were translated, and from what. The note gets written from memory, and it gets written vaguely.

The fix is workflow discipline, and it is worth choosing tools that impose it. This is the stage where Leo is built to help. Transcription runs against the page image displayed beside the text, the base ATR-1 transcription is preserved as its own tab, and every downstream operation — Correct, Modernize, Interpolate, Translate — writes to a new tab rather than overwriting it. The machine reading, your corrected reading, and any translation stay separately identifiable at the moment you sit down to describe them. Per-document metadata fields (Archive, Collection, Box, Folder, Identifier) hold the citation coordinates with the images, and TEI XML export carries the text out in a form your uncertainty encoding can live in. Because ATR-1 is a ready zero-shot model for Latin-script material — whatever the language on the page, English wills or French notarial minutes or Dutch registers alike — there is no training regime to describe in your methods note, which shortens the note to the elements that actually matter: model, date, scope, verification. Transcription and translation stay distinct operations, which is exactly the distinction your editorial note needs to draw. What Leo does not do is verify for you; the checking against the image is still yours, and it is still the part reviewers care about most.

What reviewers actually ask

Six questions, in roughly the order they arrive:

  • Which pages were machine-transcribed, and which were not? Answer by series or shelfmark, not in general terms.
  • Did you check the machine output against the images? All pages or a sample; if a sample, how it was drawn.
  • What do your brackets mean? The key, stated once.
  • Where does the machine's reading end and your editorial judgement begin? Expansions, emendations, modernized spelling, translation — each attributed.
  • Could someone repeat this? Platform, model, date, image conditions.
  • If you cite an accuracy figure, under what conditions was it obtained? Corpus, sample size, ground truth, who measured it.

None of these is hostile. They are the ordinary reproducibility questions that any method attracts, and the reason they feel sharper here is that the field has not yet settled its conventions, so reviewers are calibrating case by case. A researcher who volunteers all six in a hundred-word note has, in practice, ended the conversation before it starts. If you want the wider picture of how this stage sits alongside capture, organization, and analysis, it is worth reading the end-to-end historical research workflow as a single chain, since decisions made at capture constrain what you can honestly claim at citation.

The deeper point is one documentary editors have understood for a century: a transcription is not a neutral copy but an editorial claim about what a document says. Every mark you make — an expansion, a bracket, a decision that a crossed-out word was deleted rather than underscored — is an argument, and the note on transcription is where you show your reasoning. Machine assistance adds a line to that account. It does not change what the account is for. The editions that hold up are the ones whose authors wrote the note as though a stranger would one day open the same box, unfold the same leaf, and check.

Frequently Asked Questions

How do I cite AI transcription in a history article or monograph?

Cite the manuscript, not the machine. The archival document remains the evidentiary source and takes the full citation — repository, collection, series, box, folio — exactly as it would if you had transcribed it by hand. The recognition system is disclosed separately as a method: in a methods or sources paragraph, a "Note on transcription" in front matter, or a footnote at first use. That disclosure should name the platform and model or version, the date the pass was run, the images and pages involved, whether any training was done, how much human verification was performed against the original image, and any correction, normalization, or translation applied afterwards.

Can I use the Chicago, MLA, or APA formats for citing ChatGPT to cite HTR output?

No — those formats were written for text a model composed, and they transfer badly to text a model read off a page you can still check. Chicago treats the tool as a stand-in author, MLA routes generated material through its core-element template, and APA treats the model as a software work. All three are answering the same problem: generated text has no prior existence and is not retrievable. Handwritten text recognition produces a reading of an artefact that already exists in an archive and will still be there for a reader to verify, so the manuscript keeps the citation and the system is disclosed as method.

What should a method note about AI transcription include?

Nine elements, stated where they apply: the platform and recognition engine, distinguished if they differ; the model identifier and version, or the run date; image conditions, including whether you used your own photographs, repository surrogates, or microfilm, and at what resolution; the scope, by series or shelfmark and page count; any training and ground-truth set, or a statement that a ready model was used with no training step; the extent and method of human verification against the images; evaluation figures with their conditions attached; downstream processing such as correction, modernization, expansion, or translation; and a key explaining what your brackets mean. Roughly a hundred words is enough.

Do I still need to disclose AI transcription if I proofread every page?

Yes. Full verification is a fact you disclose, not a reason to omit disclosure — and it is usually the strongest sentence in the note. No major integrity body or publisher policy creates an exemption for fully proofread work; COPE asks that authors state which tool was used and how, in Materials and Methods or a similar section, and ICMJE requires disclosure at submission plus description in the manuscript. Publisher guidance generally speaks of generative or AI-assisted technologies rather than recognition specifically, so the safest course is to describe what you ran and what you checked, in one place, and cross-refer to it.

Can I report a vendor's published accuracy benchmark as my project's accuracy?

No. Character error rate is an evaluation figure produced under specific conditions, not a property of a tool. It depends on the hand, the script, image quality, segmentation, the quality of the ground truth itself, and the sample drawn. Report a published benchmark as a benchmark, naming its corpus, and never as your own result. If you want a number that means something for your material, measure it on your own pages: draw a randomized sample, key it by hand as reference, score it honestly, and state the sample size and composition alongside the figure. A short, conditioned local figure carries more weight with a reviewer than any headline.

Share this article