A Paleography Guide for Historians: How to Approach a Hand You Have Never Seen Before
How historians can approach an unfamiliar hand by establishing context, diagnosing the script family, building a letterform key, and fixing transcription conventions before reading, including where machine help fits.
Leo Team
August 6, 2026

Contents
This is a working paleography guide for historians facing a script they cannot yet read: a sequence for establishing context, diagnosing the hand, building a letterform key, and settling your transcription convention before you type. It matters because the alternative — reading cold and guessing from modern word shapes — is how confident misreadings enter published work and stay there.
When you open a document in a hand you cannot read, do not start at the first word. Start by establishing what the document is: its archive, its date, its genre, its language. Then form a provisional hypothesis about the script family, build a letterform key from words you can already identify with confidence — names, dates, place-names, recurring formulae — and only then read continuously, testing each doubtful graph against the key and the image. Paleography, properly understood, is not letter-by-letter decipherment. It is a method for dating, localising and interpreting written evidence, and the reading falls out of the method.
What follows is that sequence, with notes on where it changes by language, and where machine transcription can and cannot be trusted to help.
Step 1: Establish context before you read a word
Every piece of context you gather narrows the space of plausible readings. Before you attempt a doubtful word, you should know:
- Provenance. Which archive, which series, which collection. Record series tend to be internally consistent in hand, formula and abbreviation habit.
- Date and place. Even an approximate date rules out whole script families and rules in specific abbreviation systems and calendar conventions.
- Genre. A will, a notarial minute, a chancery warrant, a parish register and a merchant's ledger each have a predictable skeleton. Knowing the genre gives you the opening formula, the body, and the attestation clause before you have read them.
- Language and orthography. Not the same question as the script. More on this below.
The Archives nationales teaches its paleography course precisely this way — reading and transcribing documents while restoring them to their diplomatic, historical and institutional environment, because the institutional form of a record is evidence about its text. Codicology (the physical object: support, quires, ruling, layout) and diplomatics (documentary form, function, authenticity) sit either side of paleography, and in practice you use all three at once. The physical page tests the object; the diplomatic structure tests what kind of record it is; the letterforms test the reading.
This is a reasoning step, not a claim that scripts respect fixed dates and national borders. It means only that you should approach the page with a hypothesis rather than a blank mind.
Step 2: Diagnose the script family — provisionally
Treat the script label as a hypothesis to test, not a box to file the document in. Hands overlap, mix and hybridise. What you are looking for is a set of diagnostic features:
- Density and angle. Compressed and angular, or open and rounded? Upright or sloped?
- Minims. How does the scribe handle the vertical strokes in i n m u? Are they separated, joined, or slanted into a continuous fence?
- Ascenders and descenders. Looped? Plain? Which letters carry them?
- Single- or double-story a.
- Position-dependent forms. Does the same letter change shape initially, medially and finally?
These features do real diagnostic work. HMML's tutorial identifies Gothic Cursiva by single-story a, loops on the ascenders of b h k l, and descenders on f and the long s — and notes in the same breath that the minim-rich letters, slanting and joined to their neighbours, are exactly what makes the script hard. The University of Zurich describes Bastarda as combining features of textura and textualis while tending toward cursive, a useful reminder that the named categories are descriptive conveniences layered over a continuum.
For English material, secretary hand dominated the early modern period in England, Wales, Ireland and colonial America, frequently mixed with italic in headings, names or Latin passages within an otherwise secretary document. Some older forms persisted in specific record types as late as the 1850s. If you are working in English archives, secretary is your default first hypothesis for 1500–1700 — but a mixed hand is common enough that you should expect two letterform systems on one page. Our beginner's guide to secretary hand works through its distinctive graphs in more detail, and the broader early modern paleography guide covers the hands that surround it.
Note what the diagnosis buys you: not a reading, but a prediction. If the hand is Cursiva, you expect f and long s to be confusable. If it is secretary, you expect a particular e, a two-stroke r, and a c that looks nothing like yours.
Step 3: Build a letterform key from secure words
This is the step most often skipped, and skipping it is what produces confirmation-biased readings. The principle is local calibration: you are learning this scribe's variant of e, c, r, long s, n and u, not importing a printed alphabet chart and hoping.
Find words you can identify with near-certainty:
- A person's name that appears in the catalogue entry or elsewhere in the file
- A date, especially a written-out month or regnal year
- A place-name you already know from the provenance
- A recurring heading or field label in a register
- A formulaic phrase you can predict from the genre — In the name of God Amen, Sachez tous, Sepan cuantos
From those words, extract every letter you can. Record each graph in initial, medial and final position, because the same letter frequently changes shape by position. Note alternative forms where the scribe has more than one. Then — and this is the discipline — look for a second independent occurrence before you promote a form into the key. One instance is a guess; two is evidence.
BYU's practical advice is worth following exactly: concentrate on the lowercase consonants rather than the vowels, use ascenders and descenders as anchors, and interchange plausible graphs until the word makes sense. Their examples of what is routinely interchangeable — i / j / y, u / v, and the visual confusions between n and u, and among g y p x — are the specific substitutions to try when a word resists you. And their rule for the long s is worth committing to memory: if it looks like our f but doesn't make sense, it is their s.
Half the battle, as the same guide puts it, is recognising when the scribe has abbreviated at all.
Step 4: Read the abbreviations as a system, not as noise
Early modern abbreviation is systematic, but the system is local to language, period and scribe. The two ancient methods are contraction (omission of medial letters) and suspension (omission of terminal letters), and Cambridge's English Handwriting 1500–1700 conventions page documents the early modern repertoire built on top of them:
- Superscripts, usually contractions: w<sup>ch</sup> for which, yo<sup>r</sup> for your, ma<sup>ty</sup> for maiesty, parliam<sup>t</sup> for parliament.
- Brevigraphs: the ampersand for and or Latin et; the terminal -es graph, which may call for es, is, ys or bare s depending on the scribe's own spelling habits; the sur- graph resembling ß, which can stand for sur, ser, sir, sar or sor.
- Tildes, most often for an omitted m or n, but also for the -io- in -tion / -cion endings, and for vowel-t-vowel sequences in Latinate words.
Expansions must be consistent with the individual scribe's orthography. If the scribe prefers vocalic w, expand yo<sup>r</sup> as yowr. If the scribe writes -cion, expand reuocacõn as reuocacion, not revocation. And some tildes are pure habit — Cambridge notes that a tilde over cõme often requires no expansion at all. A guide to the main types of scribal shorthand and when to preserve or expand them is worth keeping open alongside the document.
Step 5: Decide your transcription convention before you type
Transcription is not an edition. Decide, before you start, whether you are producing:
- Diplomatic — copying what is on the page as it stands, superscripts and all
- Semi-diplomatic — Cambridge's own teaching standard, expanding contractions in the interest of readability while preserving orthography
- Normalised — modernised for a reading audience, and no longer a substitute for the source
Cambridge advises against silently regularising capitals, i/j, u/v or punctuation. Whatever you decide, make the interventions visible. The TEI Guidelines exist precisely for this: `abbr`/`expan` for abbreviation and expansion, `supplied` for editorial supply, `unclear` and `gap` for uncertainty and loss, `choice`/`sic`/`corr` for competing readings, and `facs` to bind the transcription to the image. Square brackets are one editorial convention among several; TEI markup is a different thing from silently printing an expansion, and it is the version your future self can audit.
Then verify. Recheck every occurrence of a claimed graph, every expansion, every word boundary, every recurring formula. The test to hold yourself to: a plausible word that requires a letterform the scribe never otherwise makes is a weaker reading than an awkward word supported by the image and the key. Leave unresolved readings unresolved.
The vernacular axis: same method, different vocabulary
The evidence-gathering method above is language-agnostic. Almost nothing else is. Latin script is a writing system; the language on the page may be English, French, German, Dutch, Spanish or Italian, and each brings its own hands, formulae, orthography and abbreviation habits.
- French. Fourteenth- to eighteenth-century documents, taught by the Archives nationales in their institutional context. Notarial and administrative formulae are the anchors.
- German. Kurrent and its descendants need a German-specific alphabet chart and a German vocabulary key. Do not attempt them with an English secretary mapping.
- Dutch. Dutch administrative hands and historical orthography, on their own terms.
- Spanish. The Spanish Paleography Tool supplies digitised manuscripts, typed transcriptions and sample alphabets for early-modern Cortesana, Procesal and Humanística — three distinct systems, not one "Spanish handwriting."
- Italian. The Newberry's T-PEN handbook traces vernacular production from gothic littera textualis into the formal cancelleresca and the commercial mercantesca, with hybrids throughout — a direct demonstration that script history and language history travel together.
The practical consequence: your letterform key does not transfer across languages, and neither does your formulaic phrasebook. Build a new one each time. Our overview of which languages and scripts HTR can actually read makes the same distinction from the machine side, where the confusion is even more common.
Where machine transcription fits in this sequence
The defensible position today is human-led, machine-assisted transcription. Steps 1, 2, 3 and 5 — context, diagnosis, the key, verification — remain your evidentiary control. What a machine can do is the labour of step 4: produce a literal first pass over hundreds of pages that you then check against the image, so your paleographic effort goes into the hard words rather than the easy ones.
The failure mode to guard against is specific. A fluent transcription is not a safer transcription. An evaluation of twelve multimodal models on eighteenth-century historical print found them inserting archaic characters from the wrong period — "over-historicization" — and post-correction degrading rather than improving results. Reported figures for general models on handwriting are more encouraging in places — one controlled 50-page English benchmark from 1761–1827 found competitive character error rates without hand-specific training, while itself warning about training-data contamination and transfer to other periods and languages, and still requiring verification against the image. The risk is not that a general model is always wrong; it is that when it is wrong, it is wrong in fluent, plausible, period-appropriate prose that reads exactly like a correct reading. Garbled output announces itself. A smoothed your where the page says yo<sup>r</sup> does not.
This is where a purpose-built model earns its place. Leo's ATR-1 is a transcription model trained on images of historical documents, reading Latin-script material — whatever the language written in that alphabet — without a training step of your own. It is built to transcribe what is on the page: strikethroughs, additions, marginal notes, archaic orthography, the superscript left as a superscript rather than silently resolved into a modern word. That is the property that matters at step 4. A first pass that preserves w<sup>ch</sup> is a first pass you can still make an editorial decision about; one that has already decided for you has destroyed the evidence. Errors that do get through tend to be the recoverable kind — a wrong character or word, checkable against the image displayed beside the text — rather than a confident invention. And because the base transcription sits in its own tab, anything else you want to run over it (a translation into a modern language, a glossary of difficult terms) writes to a new tab and leaves the source reading intact. Transcription and translation stay two separate jobs, which is how they should stay.
None of this replaces the key. It replaces the typing.
What the method actually gives you
The sequence — context, provisional diagnosis, key, transcription convention, verification — is a transparent working method rather than a doctrine handed down from one authority. Different historians order it differently. What it encodes is the discipline underneath: never let a reading rest on a single unverified graph, and never let readability substitute for evidence.
The compounding return is real. Two hours spent building a proper letterform key on the first three folios of a 200-page volume will save you weeks, because the rest of the volume is usually the same scribe. Two hours spent guessing at the first three folios will cost you the same weeks, later and more painfully, when a name you published turns out to have been a different name all along. The hand you have never seen before becomes, after a few hundred lines, a hand you know — and that is the actual skill, the one no tool hands you and no tool takes away.
Frequently Asked Questions
How do you use a paleography guide to read a hand you have never seen before?
Work in sequence rather than starting at the first word. First establish context: archive, series, date, place, genre, and the language as distinct from the script. Second, form a provisional hypothesis about the script family from diagnostic features — density and angle, minim treatment, ascenders and descenders, single- or double-story a, position-dependent forms. Third, build a letterform key from words you can already identify with confidence. Fourth, read the abbreviations as a system. Fifth, fix your transcription convention before typing, then verify every claimed graph against the image.
What is a letterform key and how do you build one?
A letterform key is a record of one particular scribe's graphs, built from words in the document you can identify with near-certainty — a name from the catalogue entry, a written-out date or regnal year, a place-name from the provenance, a recurring register heading, or a formulaic opening such as In the name of God Amen. Extract every letter from those secure words, recording each in initial, medial and final position, and note alternative forms. Require a second independent occurrence before promoting a form into the key: one instance is a guess, two is evidence.
What script should I expect in English documents from 1500 to 1700?
Secretary hand is the default first hypothesis for English material in that period, and it dominated in England, Wales, Ireland and colonial America. Expect it frequently mixed with italic — in headings, names, or Latin passages inside an otherwise secretary document — so plan for two letterform systems on one page. Some older forms persisted in specific record types as late as the 1850s. Secretary brings a particular e, a two-stroke r, and a c that looks nothing like the modern letter, along with the usual long s and f confusion.
What is the difference between diplomatic, semi-diplomatic and normalised transcription?
Diplomatic transcription copies what is on the page as it stands, superscripts and all. Semi-diplomatic expands contractions for readability while preserving the scribe's orthography; this is Cambridge's teaching standard. Normalised transcription modernises the text for a reading audience and is no longer a substitute for the source. Decide which you are producing before you type, avoid silently regularising capitals, i/j, u/v or punctuation, and make every editorial intervention visible. TEI markup gives you elements for abbreviation and expansion, editorial supply, uncertainty, loss, competing readings and binding the text to the image.
Why is a fluent AI transcription more dangerous than a garbled one?
Because garbled output announces itself and fluent output does not. A smoothed your where the page reads yo<sup>r</sup> looks exactly like a correct reading, so the error survives verification unless you check against the image. An evaluation of twelve multimodal models on eighteenth-century historical print found them inserting archaic characters from the wrong period — over-historicization — with post-correction degrading rather than improving results. The defensible position is human-led, machine-assisted work: let a purpose-built model produce a literal first pass that preserves superscripts and archaic orthography, and keep context, diagnosis, the key and verification as your evidentiary control.