OCR for Old Books: Why Historical Print Breaks Standard OCR, and What to Use Instead

Why standard OCR misreads old books even when a person reads them easily, what newer whole-line models change, and how to choose a model by the kind of book. It also covers scanning and photographing old books, and getting searchable text out.

Leo Team

October 4, 2026

OCR for Old Books: Why Historical Print Breaks Standard OCR, and What to Use Instead
Contents

This is a guide to OCR for old books: why standard OCR misreads historical print that a person reads without trouble, what the newer whole-line models change, how to choose a model by the kind of book in front of you, and how to scan or photograph it and get usable, searchable text out.

Run a clean twentieth-century novel through any OCR engine and the text comes back almost perfect. Run a book printed in 1700 through the same engine and you may get "fame" where the page says "same", words broken at every ligature, and the author's spelling "corrected" into modern English. The type is perfectly legible to you. The trouble is that standard OCR was built for a different kind of print.

This guide is about printed books and pamphlets. For handwriting, start with our guide to handwritten text recognition, and for how the two technologies differ, see the difference between OCR and HTR.

Why standard OCR breaks on old books

Conventional OCR engines were designed around modern, regular type. They split the page into blocks, lines and characters, match each character against the letterforms they know, and then check each word against a dictionary of the language. On a modern book every one of those steps works. On an old book each one can go wrong.

The effect has been measured for a long time. A 2009 British Library study found that fully automated OCR of anything printed before 1900 would be "fortunate to exceed 85%" character accuracy, and its 19th-century newspapers came out at 78% word accuracy (Tanner, Muñoz & Ros, 2009). Words beginning with a capital, which is where names and places sit, came out lower still, at 63.4%. Those are the words you are most likely to search for.

The causes are specific.

The long s read as f

Until about the end of the eighteenth century, printers used the long s (ſ) everywhere except at the end of a word. In roman type it looks almost exactly like an f without the full crossbar, and an engine trained on modern fonts has no ſ to match it to. So "ſame" comes out as "fame" and "roſe" as "rofe". In one study of an eighteenth-century collection, words containing "f" made up 11% of a hand-keyed version of the texts but 33.7% of the OCR'd version (Hill & Hengchen, 2019).

Ligatures

Hand-press printers cast common letter pairs as single pieces of type: ct, ff, ffl, and several with the long s. An engine that does not know these shapes splits them wrongly or drops them. The same study found that words containing ligatures with a long s, or ct, ff and ffl, were significantly more likely to be misread.

Blackletter and mixed type

Blackletter faces, including Fraktur, which German printers used well into the twentieth century, have letters that look alike to a machine trained on roman type: n and u, f and ſ, c and e. Many old books also mix blackletter, roman and italic on one page, sometimes in one sentence. Our guide to Fraktur OCR covers this type in detail.

The physical page

Old books were printed by hand on rag paper, and the page shows it: uneven inking, worn or broken type, wavering lines, foxing, and show-through from the other side of a thin leaf. A character classifier treats all of it as evidence, so a smudge becomes a letter and a faint letter becomes nothing. Hill and Hengchen list smudges, damage and shine-through among the physical causes of OCR error.

Period spelling "corrected" by a modern dictionary

The dictionary step that rescues a modern page works against an old one. Early printers spelled "vnto", "haue" and "bloud", and an engine that checks its guesses against a modern word list pulls those spellings towards something it recognises. Tanner and colleagues suggest this is one reason why the significant words on a page, which are often longer and missing from dictionaries, are harder to OCR accurately. Abbreviation marks inherited from scribes, such as the macron over a vowel standing for an omitted m or n, are often missing from the engine's character set altogether, so they are dropped or silently expanded.

Tables and maps

Even on clean nineteenth-century print, layout matters. A test published by UC Berkeley Library in December 2024 found that ABBYY FineReader excelled at paragraphs of printed text, but picked up only about 25% of the text on maps and 80% of the data from tables in the same documents (Berkeley Library, 2024). Check tables and maps by hand, whatever tool you use.

One more warning applies to all of the above: the confidence score an OCR engine reports is not a reliable guide. In Hill and Hengchen's study, the OCR software overestimated its own accuracy.

What newer whole-line models change

The newer generation of models reads differently. Instead of matching characters one at a time and correcting them afterwards, it reads a whole line or page in context, so the surrounding word tells it that the tall letter in "ſame" is an s, not an f.

The difference on old print is large. In a 2026 study comparing OCR tools on archival material, the old engines got roughly a fifth of the characters wrong on worn seventeenth-century English print, while the best new tools got about one in forty wrong (Clifford et al., 2026). The same study found that on clean, single-column nineteenth-century print the modern tools were effectively tied, so the choice there comes down to speed and cost. Early-modern print was about four times harder for the same tools than Victorian print. The authors conclude that the difficulty lies in the type and spelling, not the layout.

These models fail differently, though. Their mistakes are no longer garbled lines you can see and discard. They are fluent, plausible misreadings that read naturally, which is why the same authors call them the dangerous kind. The commonest is quiet modernisation: the study records "bloud" silently becoming "blood". Told to keep the original spelling, a model became more faithful, which shows where the default pull lies. For a reader, a modernised word is a small slip. For anyone quoting the text, studying spelling, or searching for a variant form, it is lost evidence. Our guides to fluent errors and historical language normalization explain why the faithful text and a modernised one should be kept apart.

Choose by the kind of book

The right tool depends on what is printed on the page and what you need the text for. Old books fall into a few clear groups.

A clean 19th- or 20th-century book, or a typescript

A well-printed book from the nineteenth or twentieth century, or a clean typescript, is the easy case. The type is regular, the spelling is close to modern, and almost any current model reads it well. Cost and speed decide.

In Leo, this is where a small, inexpensive model such as GPT-6 Luna fits. It costs 0.02 credits an image, so 1,000 images come to about 20 credits. On public plans a credit costs $0.12–$0.20, which puts a 1,000-image book at a few dollars. In our tests it read clean, single-column books and typescripts as well as the alternatives. Spot-check a few images from the start, middle and end, and give footnotes and small type a closer look.

Early-modern print, roughly 1500 to 1800

Here the job changes. The long s, ligatures, macrons, u and v used interchangeably, and spelling that was not yet fixed all appear on the page, and the transcription is only useful if it keeps them. A cheap model is not the right tool: in our tests, small models modernised early-modern spelling.

This is the work Leo's own model, ATR-1, is built for. It is specially tuned for historical documents, printed and handwritten, and trained to transcribe what is on the page rather than normalise it. The long s stays a long s, the macron is not silently resolved, and "bloud" stays "bloud". If you will quote the text, study its spelling, or prepare an edition, start here.

Blackletter and Fraktur

Blackletter and Fraktur are a job for ATR-1, or a stronger general model as a second reading, not for a cheap one. In our tests, small models misread names on Fraktur pages. Judge the result on a sample of your own pages before you commit a whole volume.

Newspapers and other multi-column pages

Old newspapers add a layout problem to the type problem: several narrow columns, headlines, advertisements and running heads, which a tool has to put in the right order. Even when every word is read correctly, the text can come out in the wrong order. Some models also stop partway down a dense page without saying so: in our tests, small models sometimes read only a fraction of a crowded newspaper page and returned what looked like a finished transcription. Check the end of each page against the image, and test any tool on a few of your own pages before running a whole run of issues.

Whatever the book, the surest guide is a small trial on your own material. Our method for how to test AI transcription on your own documents takes an afternoon.

Scanning and photographing an old book for OCR

A good image does more for OCR than any setting. The general advice on photographing archival documents applies, and books add a few rules of their own:

  • Never force the spine. Support the book on a cradle or folded towels so it opens only as far as it wants to. Photograph each page separately, tilting the camera to stay square to the page, rather than pressing the book flat.
  • Fight the curve at the gutter. Text that bends into the spine distorts the letters. Hold the page gently at its outer edge with a strip of card, or lay a sheet of clean glass over it if the binding allows.
  • Use colour or greyscale, never black and white. Black-and-white capture throws away faint strokes and the difference between ink and show-through. For resolution, 300 dpi suits normal type and 600 dpi suits small type. Our guide to what DPI to scan old documents explains why.
  • Slip black card behind a thin leaf. It cuts the show-through from the other side.
  • Keep the margins. Include page numbers, running heads, catchwords and any marginal notes in the frame, then decide later what to keep.
  • Use even light and no flash. A flash makes glare on glossy or cockled paper and hides the letters under it.
  • Name files in page order, for example vol1-p023.jpg, so the transcription follows the book

Many old books are already online as scans in digital libraries, often with a text layer made by conventional OCR. Download the images, not the text: that layer usually comes from the engines described above, with the long s read as f and the old spelling "corrected". In Leo you can upload such a PDF and choose to convert each of its pages into an image, which Leo recommends.

From image to searchable text: the output

Decide what the text is for before you start, because it decides how faithful it must be.

  • For reading or a quick search, a clean modern-spelling text is enough
  • For quoting, editing or language study, you need a faithful transcription that keeps the long s, the abbreviations and the spelling as printed
  • Often you want both. Keep the faithful transcription as the record, and make a modern-spelling version alongside it for reading and searching. Never let the modern version replace the faithful one.

Then check names, dates, numbers and anything you plan to quote against the image. A fluent model's errors look right, so they have to be checked deliberately, not spotted in passing.

Where Leo fits

Leo is a transcription workspace for historical documents, printed and handwritten, in any language written in the Latin alphabet. For old books, it puts the choice of model in one place.

In the transcription window, one menu holds Leo's own model, ATR-1, as the default, alongside the major general models from OpenAI, Google and Anthropic and small, inexpensive models for clean print. Prices are per image and depend on the model. ATR-1 is always 1 credit an image, and GPT-6 Luna is 0.02 credits. A job rounds up to a whole credit, so the smallest job is 1 credit. See plans and credit packs.

ATR-1 is the model to use for early-modern print and anything you will quote. It is trained to transcribe what is on the page rather than tidy it, so the long s, the macron and the period spelling survive into the transcription, not quietly modernised the way general-purpose AI models tend to do. It reads the image at high resolution, where general models downsample it, and needs no training on your book. Built-in failure detection hides output that looks wrong, retries it, and refunds the credit if it cannot succeed. The errors that do get through are the recoverable kind: a wrong character or word that you check against the image shown beside the text. Every correction you make feeds back into training, so later readings improve.

The rest of the workflow suits a book or a run of volumes:

  • Upload scans, phone photographs or a whole PDF, one image per page
  • Compare readings. Run a second model over the same image and its reading lands in its own tab, so you can set two readings side by side. The ATR-1 transcription is never overwritten.
  • Keep layers separate. The Modernize Transformation writes a modern-spelling version into its own tab, beside the faithful transcription.
  • Search every transcription at once, with fuzzy search to catch variant spellings
  • Export to Word, PDF, HTML or TEI, with line breaks kept if you want them, or share a read-only link that needs no account

Leo has a free plan, so you can run a few images from your own book through ATR-1 and a cheaper model, and see which one the book needs.

Frequently Asked Questions

Why does OCR read the long s as f?

The long s (ſ) looks almost exactly like an f in roman type, missing only part of the crossbar, and it was used everywhere except at the end of a word until about 1800. Engines trained on modern fonts have no ſ to match it to, so they choose f. In one study of an eighteenth-century collection, words with "f" were 11% of a hand-keyed version of the texts but 33.7% of the OCR'd one. Models that read the whole line in context avoid most of these errors, and Leo's ATR-1, built for historical print, keeps the long s as printed.

What is the best way to OCR an old book?

Choose the model by the kind of book. A clean nineteenth- or twentieth-century book or a typescript is easy, and a small, inexpensive model reads it well: in Leo, GPT-6 Luna at 0.02 credits an image. Early-modern print, blackletter and anything you will quote need a model that keeps the spelling on the page, such as Leo's ATR-1. Photograph or scan in colour or greyscale, support the spine, and test on a few images before running the whole book.

Should OCR keep the original spelling of an old book?

Yes, for the record. Old spelling, the long s, u and v used interchangeably and abbreviation marks are part of the text, and a quietly modernised word is lost evidence for anyone who quotes or studies it. If you want an easier text for reading or searching, make a modern-spelling version separately and label it. In Leo, ATR-1 keeps the spelling as printed, and the Modernize Transformation writes a modern version into its own tab.

How do I make an old book searchable?

Photograph or scan every page, transcribe the images with a model suited to the type, and check names and key passages against the images. Keep a faithful transcription, and add a modern-spelling version if you want searches to find modern forms. In Leo you can then search every transcription at once, with fuzzy search to catch variant spellings, or export the text to Word, PDF, HTML or TEI.

How much does it cost to OCR an old book?

In Leo it depends on the model, priced per image. A clean, legible later book on GPT-6 Luna costs 0.02 credits an image, so 1,000 images come to about 20 credits, a few dollars at public prices of $0.12–$0.20 a credit. ATR-1, for early-modern print and anything you will quote, is 1 credit an image. A job rounds up to a whole credit, and there is a free plan to try it on your own book.

Share this article