Can AI Read Faded Documents? Faded Ink, Foxing, and Show-Through, Triaged

Triage guide for archivists on AI reading of faded documents, sorting capture, recognition, and conservation problems for iron-gall ink, foxing, and show-through.

Leo Team

August 6, 2026

Can AI Read Faded Documents? Faded Ink, Foxing, and Show-Through, Triaged
Contents

Can AI read faded documents? Often yes — provided the ink signal still exists in the image you captured, which makes this as much a photography and conservation question as a recognition one. This piece is a triage guide for archivists, librarians, and collections staff: how to tell a capture problem from a recognition problem from a conservation problem, and what each one actually requires.

AI transcription can read many damaged pages — faded iron-gall ink, foxed paper, light show-through from the verso — when the ink signal is present in the capture. What no recognition system can do is recover text that has physically gone: flaked ink, holes, losses, corroded-out strokes. So the practical question is not "which AI reads damaged documents" but "is this a capture problem, a recognition problem, or a conservation problem?" Each has a different remedy, and only one of them is solved by software.

What follows works through that triage in order: what the damage does to the pixels, what capture can fix, what preprocessing can and cannot fix, where recognition sits in the chain, and how to decide, page by page, between re-photographing, accepting with gaps, and referring the item to a specialist.

What damage does to an image, condition by condition

Recognition systems consume pixels and layout. They do not know that a mark is a stain. The useful way to think about damage is to ask what it does to the difference between ink and support — and whether that difference still exists anywhere in the captured signal.

Iron-gall ink corrosion

Iron-gall ink is an iron/tannin/gum system, and the reason so much of the European archival record is fading is chemical rather than optical. Excess iron and the ink's acidity drive oxidation and acid hydrolysis, with iron catalysing cellulose-chain scission. IFLA's guidance on iron gall ink lists the signs: a gradual shift from black to brown, brown discolouration and embrittlement of the support, friable or cracked inked lines, loss of inked areas, and spread of the reactions into adjacent material. A review of iron-gall degradation mechanisms in npj Heritage Science frames it plainly: this is heritage at risk of total loss of the support itself.

What a camera receives, then, is not "difficult handwriting." It is broken strokes, halos where iron has migrated into previously uninked paper, and contrast that varies across the page rather than uniformly. A recognition model faced with a stroke that is 60% present has to guess the remaining 40%. That is exactly where the difference between a careful transcription engine and a fluent one starts to matter.

Foxing

Foxing is a description, not a diagnosis: scattered spots, commonly reddish-brown but ranging from yellow to black, distinct from visible mould colonies though the two can co-occur. The AIC Book and Paper Group's chapter on foxing is candid that after nearly sixty years of investigation the causes remain divided — fungal activity and metal-induced oxidative degradation both have evidential support, and many conservators believe a combination is responsible. Foxed paper is more acidic and has lower tensile strength.

For imaging purposes, the relevant fact is that foxing produces local background variation at roughly the scale of letterforms. Spots can read as foreground marks; more insidiously, they break the assumption that the page background is stable, which is what disrupts segmentation before recognition even starts. No defensible published figure isolates foxing's effect on character error rate, and one should be treated with suspicion if offered: that step is a reasonable image-processing inference, not a measured law.

Show-through versus bleed-through

These get conflated constantly, and the distinction determines whether the problem is solvable.

Show-through is optical: verso marks visible through a translucent sheet. The ink is still on the other side. Bleed-through is physical: ink has penetrated or migrated into the support and is genuinely present on this side of the paper.

That difference matters because show-through can sometimes be modelled and separated computationally when both sides are captured and registered — a 2024 restoration study works from two scanned images, front and back, for exactly this purpose. Irreversible migration cannot simply be un-seen. Neither result generalizes to a promised success rate, and the method fails when only one side is available, when recto and verso cannot be registered, or when the overlap is heavy.

The operational implication is cheap and worth adopting as policy: capture the verso even for blank-looking backs of translucent leaves. It costs one exposure and preserves an option you cannot reconstruct later.

Water, mould, and structural damage

Water moves soluble material and evaporates into tide lines, changing both paper dimensions and optical background; dampness promotes mould and can intensify iron-gall corrosion. Mould stains and chemically damages cellulose. Tight gutters, curvature, skew and cockling bend baselines, cast shadows and hide text — layout analysis can fail on geometry alone, before a single character is read. Tears, insect losses and embrittlement remove information outright. Software can interpolate a plausible shape into a hole. It cannot establish what was there.

Capture problems versus recognition problems

Before evaluating any tool, sort the page.

  • Capture-side problems: focus, exposure, glare, insufficient sampling, colour and tonal error, gutter occlusion, page curvature, ordinary-RGB invisibility, unrecorded verso information. These are fixable by re-photographing, using a cradle or a safer opening angle, controlling the light, shooting both sides, or moving to a different imaging modality entirely.
  • Recognition-side problems: the evidence is in the file, but the pipeline mishandles it. Remedies include layout analysis, deskew and dewarp, background normalization, denoising, contrast enhancement, adaptive binarization, two-sided separation, or simply a recognition model better matched to the hand and period.
  • Neither: physical loss, severe corrosion, truly absent spectral signal. This is a conservation and paleography question, and throughput should yield to it.

Getting the sort wrong is expensive in both directions: months spent trialling recognition engines on an image that a better photograph would have solved in an afternoon, or repeated handling of a fragile volume in pursuit of information that is chemically no longer there. Much of this is decided upstream, at the point where a digitization workflow sets capture standards for a whole series — which is why capture is the highest-leverage layer in the chain.

What the imaging standards do and do not promise

Metamorfoze Preservation Imaging Guidelines v2.0 (2025) sets a normal sampling rate of 300 ppi, a minimum detail level of 5 line-pairs/mm, MTF10 performance of 5 lp/mm and sampling efficiency of at least 85%; originals smaller than A5 may go above 300 and up to 600 ppi with written consent, and Extra Light permits up to 4% geometric distortion for textual information. FADGI's third edition is explicitly informative rather than prescriptive, notes that resolutions above 400 ppi may be appropriate with 800 ppi or more for very fine detail, and stresses that conformance requires a managed QA programme.

Both preserve visible information and measurable image quality. Neither supplies a transcription-grade ppi, and neither proves an OCR or HTR outcome. There is no standardized published link from FADGI/Metamorfoze/ISO measurements to downstream character error rate for vernacular historical material. Meeting the standard is necessary discipline; it is not a guarantee that a machine will read the page.

When the information is spectrally present but invisible

If the ink signal may still exist outside ordinary RGB, capture can change the problem entirely. Multispectral imaging records many wavelengths and combines them to separate the ink's spectral response from the support and from overlying material. The Archimedes Palimpsest imaging project is the standing demonstration: imagers separated the Archimedes ink's signature from the parchment beneath and the prayer book above it, recovering text and diagrams invisible or nearly invisible under RGB light.

That is proof of possibility, not a promise for every faded parish register. IR, UV fluorescence, transmitted light, raking light and RTI exploit different physical properties — wavelength-dependent reflectance and fluorescence, sheet transparency, surface relief, changing illumination — and their reach is object-specific. All of them require specialist cameras, filters, lighting, conservation-safe handling and expert interpretation. For most institutions these are referral routes to an imaging lab or conservation service, not a setting in a workflow. Comparative cost-benefit studies for ordinary English, French, German, Dutch, Spanish or Italian archival series are thin, and it is not well established when repeated capture and illumination risk outweighs the expected information gain.

Preprocessing: mature in concept, conditional in outcome

The standard families are well understood: deskew and crop, illumination correction, contrast normalization, denoising, local or adaptive thresholding, layout and reading-order analysis, line segmentation, dewarping.

Adaptive methods exist precisely because damaged pages defeat global ones. As the review of degraded historical document binarization sets out, global thresholding applies a single value to the whole image and assumes a relatively stable background; local methods such as Sauvola compute thresholds from local pixel neighbourhoods, which is the only way to represent uneven paper tone, stains and illumination on the same sheet.

Two cautions belong with every preprocessing decision.

  1. Thresholding can erase faint ink. The same operation that removes foxing spots can remove a corroded stroke, and the output will look cleaner while containing less evidence.
  2. Learned restoration can invent clean-looking strokes. Deep-learning restoration and learned binarization can outperform hand-designed rules on particular datasets, but remain research-stage for heterogeneous heritage damage, and can produce a visually convincing result that is historically unverified.

So: keep the RGB and any multispectral masters unprocessed, log the processing chain reversibly, and never let a cleaned derivative replace the preservation master. More preprocessing does not reliably fix a bad scan; it reliably makes a scan look fixed.

What recognition can actually do with a damaged page

Assume you have the best safe image. What reads it?

Modern-print OCR, historical-print pipelines and handwriting recognition solve genuinely different distributions, and the categories are not interchangeable. General cloud OCR is useful on clean or moderately difficult machine print, and some services expose handwriting and layout APIs — but there is no fair, independent, apples-to-apples published comparison of Tesseract, ABBYY, Google Vision and Textract on damaged vernacular heritage material, and Tesseract's own guidance warns that skew sharply reduces line-segmentation quality. Historical print needs models that know long s, ligatures, typographic abbreviation, blackletter and Fraktur, and complex layout: OCR4all and OCR-D exist as workflow ecosystems for exactly that reason. For handwriting, specialist HTR is the strongest established route — Transkribus, eScriptorium/Kraken — when a compatible model exists or you can train one, which means aligned images and ground-truth transcriptions for the hand and language in front of you.

Whatever the engine, the image sets the ceiling.

The genuine danger on damaged pages is not garbled output. It is fluent output. A generative vision-language model faced with ambiguous pixels can complete a plausible name or word using language priors rather than visual evidence — work on OCR hallucination in multimodal LLMs identifies visual degradation (blur, occlusion, low contrast) as precisely the condition under which models over-rely on linguistic priors and fail to register their own uncertainty. A study of multimodal LLMs on historical handwriting found that where models failed badly, the failure was usually text generation rather than letterform recognition, sometimes producing text entirely unrelated to the image. And while a 2025 Journal of Documentation benchmark of LLMs for handwritten text recognition reported strong figures on modern English handwriting — with results skewed towards English and declining on historical and other-language material, across a limited set of available datasets, none of them damage-controlled — those numbers describe clean modern pages and cannot be relabelled as faded-page results.

An engine that returns a wrong character, or refuses a character, hands you a recoverable error you can catch against the image. An engine that returns a confident, well-formed surname where the paper has a foxing spot hands you an error that will survive into your catalogue. This is the difference between fluent and correct that matters most on damaged material, and it is why confidence scores should be treated as a routing signal rather than proof — calibration varies by engine and domain, and confidence-based error detection weakens under exactly the severe degradation where you most want it.

Where a specialist transcription model fits

For an archive with a damaged but visible series in a Latin-script hand — English wills with tide lines, French notarial records with show-through, Dutch registers with foxing — the practical need is a model that reads the marks as they survive and does not tidy them.

This is the stage Leo occupies. Its recognition model, ATR-1, is zero-shot: no per-collection training step and no ground-truth alignment before you can see what a series looks like transcribed, which matters when you are triaging a damaged run and need a directional answer this week rather than after a training cycle. It reads Latin-script material — any language written in that alphabet, English, French, German, Dutch, Spanish, Italian and others among them, with non-Latin scripts out of scope — in both manuscript hands and print, including the early-modern typography that trips conventional OCR. On a damaged page the relevant commitment is source integrity: transcribe what is on the page, preserving strikethroughs, additions, margin notes and archaic orthography, rather than normalizing them into modern prose. Against fabrication it offers a mechanism rather than a promise. Output matching known failure patterns — a line repeating over and over, the loop any text-generating model can fall into on material unlike its training data — is hidden, retried automatically, and if it still cannot succeed the job stops and the credit is refunded. Errors that get through tend to be the recoverable kind, checked against the image displayed beside the text.

Two honest limits. Pages that mix dominant printed structure with dense handwriting — pre-printed ledger and register forms — are the known weak spot, where the model can favour the printed elements; complex tabular layouts vary. And Leo begins at upload: it does not capture, and it will not recover a signal your photograph did not record. On head-to-head manuscript accuracy the published comparison is a clean-condition one — a randomized 97-image sample of early-modern English manuscripts from the Folger Shakespeare Library, where ATR-1 scored roughly 5% character error rate against Transkribus/Text Titan I at ~13%, Claude Opus ~23.3%, Gemini 2.5 Pro ~24.8% and GPT-4.1 ~56.7% (full results). That is a manuscript benchmark, not a damage benchmark, and no vendor figure — this one included — transfers to your foxed volumes. Measure a representative sample of your own material before committing a series.

A triage boundary you can apply page by page

Damaged material rewards a decision rule more than a tool preference.

Re-photograph when the evidence is probably present but the capture is deficient: glare, shallow focus, gutter loss, low sampling, no verso image. Cheapest fix in the chain, and the one most often skipped.

Refer for special capture when the ink may still be spectrally present but is invisible in RGB — heavy fading, corroded iron-gall, palimpsest-like overwriting. Expect equipment, conservation-safe handling and expert interpretation, and expect object-specific results.

Transcribe with machine assistance and verify when the text is visible but the hand or condition is hard. Route names, dates, numbers, sums and boundary descriptions to human review as a matter of course, mark illegible spans explicitly rather than letting any system guess them, and audit a random sample of high-confidence lines as well as the flagged ones.

Use multiple human readers when the ambiguity is visible and resolvable by eye. This is the case where crowdsourced transcription genuinely earns its place — but scale it honestly. A 2025 study of collaborative HTR workflows worked from a corpus of about 100,000 manuscript pages and, with 3,926 registered volunteers plus anonymous contributors, transcribed 14,330 images over twelve months; a limited agreement test across 20 folios produced a Cohen's kappa of 0.66, described as moderate. Its HTR models trained on 490 and then 900 pages reached 6.99% and 5.56% CER, but on two sampled pages error rates ranged from 7.44% to 68.76% depending on model and page. Page-level variance, not the average, is what damaged material produces.

Accept with explicit gaps when the best safe image still cannot resolve a mark and the record is low-risk discovery material. A catalogued gap is a defensible scholarly statement. A guessed name is not.

Stop and refer to a conservator when the object is chemically unstable, structurally unsafe to open, or materially incomplete. Nothing downstream is worth further loss.

The habit worth keeping

The instinct to reach for a better model on a damaged page is usually misdirected effort. Condition determines what the image contains; the image determines what any recognition system can see; and the recognition step, however good, is bounded by both. Work in that order and the tooling question mostly answers itself.

What remains is judgement of the kind archives have always exercised: knowing which volumes deserve a second photograph, which deserve a conservator, and which can go out into the catalogue with an honest bracket where the paper has failed. Record what you did to the image, keep the master untouched, and let the gap stand where the evidence stands nowhere. A collection described accurately, holes and all, is more use to the next researcher than one that reads smoothly and cannot be trusted.

Frequently Asked Questions

Can AI read faded documents like iron-gall ink manuscripts and foxed pages?

Often yes, provided the ink signal still exists in the image you captured. AI transcription can handle faded iron-gall ink, foxed paper and light show-through from the verso when enough contrast between ink and support survives in the photograph. What no recognition system can do is recover text that has physically gone — flaked ink, holes, losses, corroded-out strokes. So the useful question is whether you have a capture problem, a recognition problem, or a conservation problem. Each has a different remedy, and only one of them is solved by software. Sort the page before you evaluate any tool.

What is the difference between show-through and bleed-through, and does it matter for transcription?

Show-through is optical: marks on the verso are visible through a translucent sheet, but the ink is still on the other side. Bleed-through is physical: ink has penetrated or migrated into the support and is genuinely present on this side of the paper. The distinction determines solvability. Show-through can sometimes be modelled and separated computationally when both sides are captured and registered; irreversible migration cannot be un-seen. The method fails when only one side is available, when recto and verso cannot be registered, or when overlap is heavy. Practical policy: capture the verso even for blank-looking backs of translucent leaves.

Does scanning at higher resolution guarantee better AI transcription of damaged pages?

No. Metamorfoze and FADGI guidelines preserve visible information and measurable image quality, but neither supplies a transcription-grade ppi, and neither proves an OCR or HTR outcome. There is no standardized published link from those measurements to downstream character error rate for vernacular historical material. Meeting the standard is necessary discipline, not a guarantee that a machine will read the page. Resolution also cannot add signal that was never recorded — glare, shallow focus, gutter occlusion and an unphotographed verso are capture faults that a higher ppi setting will not correct. Re-photographing is usually the cheapest fix in the chain.

Why is a confident AI transcription more dangerous than a garbled one on damaged documents?

Because a wrong or refused character is a recoverable error you can catch against the image, while a confident, well-formed surname invented where the paper has a foxing spot will survive into your catalogue. Generative vision-language models faced with ambiguous pixels can complete plausible words from language priors rather than visual evidence, and visual degradation — blur, occlusion, low contrast — is precisely the condition under which they over-rely on those priors and fail to register uncertainty. Confidence scores should be treated as a routing signal, not proof: calibration varies by engine, and confidence-based error detection weakens under severe degradation.

Can preprocessing or image restoration recover faded ink before transcription?

Sometimes, but with two cautions. Adaptive methods such as local thresholding exist because damaged pages defeat global ones, and illumination correction, denoising and dewarping genuinely help. But thresholding can erase faint ink — the same operation that removes foxing spots can remove a corroded stroke, leaving output that looks cleaner while containing less evidence. Learned restoration can invent clean-looking strokes that are visually convincing and historically unverified. Keep RGB and any multispectral masters unprocessed, log the processing chain reversibly, and never let a cleaned derivative replace the preservation master. More preprocessing does not reliably fix a bad scan; it makes a scan look fixed.

Share this article

© 2026 Leo Technologies Limited. All rights reserved