How to Write a Finding Aid: From Survey and Arrangement to a Published, Searchable Guide
How to write a finding aid as a working sequence from survey and arrangement through description depth, fidelity to the records, and publication as a searchable guide.
Leo Team
August 18, 2026

Contents
This is a working sequence for how to write a finding aid — survey, arrangement, level of description, the descriptive notes, fidelity to the record, and publication — with the judgment calls named where they actually fall. If you are responsible for making a collection findable, the two decisions that shape everything downstream are how deep you describe and how faithfully you record what the records say.
A finding aid is a structured guide describing the intellectual arrangement and content of a processed collection — a research road map, not an item catalogue (SAA Dictionary of Archives Terminology). Writing one runs in a fixed order: survey the whole accession, establish provenance and the best available order, choose the level or levels of description the material justifies, then write the description — reference code, title, dates, extent, creator, scope and content, arrangement, access conditions, languages and scripts, rights — and finally publish it in a form that is both readable by researchers and indexable by machines.
What follows is that sequence.
Stage one: survey before you describe
The instinct to start typing into ArchivesSpace is strong, and it is the wrong instinct. Description documents contextual relationships; you cannot document relationships you have not yet identified. Johns Hopkins' processing manual — useful institutional practice rather than universal standard — describes the survey as a pass through the entire collection to determine what is present and how it should be arranged (JHU Special Collections best practices). That means opening boxes you would rather not open.
The survey should produce four things:
- A provenance statement. Who created, accumulated, maintained, or used these records? Provenance is the relationship between records and their creator, and respect des fonds requires that records of different origins stay separate so that context of creation and custody remains legible.
- A record of the received order. Not because received order is sacred — it usually isn't — but because you cannot disclose what you changed if you did not note what arrived.
- A preservation and restrictions triage. Fragile material, mould, and anything carrying privacy or third-party rights problems needs flagging before it reaches a reading room.
- A processing plan that names the level of description you intend to reach, and why.
Two misconceptions get corrected here. The first is that arrangement can be fixed later, after description is drafted. In practice, re-arranging after describing means rewriting the description. The second is that the most helpful archive is sorted by subject. It isn't: subject sorting destroys evidence of who created records, in what functions, and how activities were carried out. DACS is direct that aggregating resources by common subject is less useful when it obscures how and by whom records were made. Subject access belongs in access points and scope notes, not in the physical or intellectual re-sorting of a fonds.
Original order deserves a more careful reading than it usually gets. Current DACS principles treat it as an intellectual and contextual construct — the creator's recordkeeping groupings and relationships — rather than as a single pristine physical sequence. The order in which boxes arrived may not be the creator's last working order, and there may be several meaningful orders. Where no usable order survives, impose the least distortive workable arrangement, document that you imposed it, and describe it honestly. Disclosure is what keeps an imposed order from becoming a silent fabrication.
Stage two: choose your level of description
ISAD(G) sets out a hierarchical framework — fonds, series, file, item — and requires that description move from the broad to the specific, identify whole–part relationships, and avoid repeating higher-level information further down (ISAD(G), 2nd edition). DACS is more permissive: a finding aid may be single-level or multilevel, may begin at any archival level, and DACS deliberately does not prescribe a correct depth.
That leaves the depth decision with you. The useful framing is cost–benefit rather than completeness. Collection- or series-level description is usually adequate for large, homogeneous, or low-risk material. File- or item-level description is justified where:
- titles do not communicate content (a folder marked "Correspondence" tells a researcher nothing);
- the material carries high research or evidential value;
- rights, privacy, or access decisions have to be made unit by unit;
- expected access benefit exceeds the processing and ongoing maintenance cost.
DACS supports this proportionality in its own minimum rules: scope and content is typically needed for large aggregations and is not required at file or item level where the title sufficiently describes the material.
More Product, Less Process is the standing strategy for backlogs, and it is not a synonym for neglect. Greene and Meissner's 2005 American Archivist article reported survey findings that 34% of responding repositories had more than half their holdings unprocessed and 60% had at least a third unprocessed — figures drawn from 100 returned surveys, which the authors themselves treated as suggestive rather than as a prevalence estimate. MPLP's answer is a comprehensible arrangement, preservation triage, restrictions control, and an overview with a container inventory, while skipping low-value refinements. The follow-up study documents adoption and reported processing rates, along with an honest caveat: coarser description can shift labour to reference and retrieval, because thinner description means pulling more boxes. That trade-off is a policy judgment, not a settled quantitative law. Make the processing level explicit in the finding aid, and document what was not done.
Stage three: the elements, and which are actually required
A working finding aid normally combines: repository and reference code; title; creator; dates; extent; a collection-level summary; administrative or biographical history; scope and content; an arrangement statement; the hierarchical series and container list; conditions governing access and reproduction; languages and scripts; rights; related materials; and description-control information.
Not all of these are equally mandatory. Two reference points are worth holding separately.
The DACS single-level minimum
Reference code, name and location of repository, title, date, extent, name of creator (if known), scope and content, conditions governing access, languages and scripts of the material, and rights statements. A multilevel DACS description applies that minimum at the top, states the whole–part relationship, and carries appropriate information down the hierarchy — with creators recorded at lower levels where they differ. Note where DACS places its weight: administrative and biographical history is an optimum element; scope and content is required.
The ISAD(G) framework
Twenty-six elements across seven areas (Identity Statement, Context, Content and Structure, Conditions of Access and Use, Allied Materials, Note, Description Control), of which only six are essential for international exchange — reference code, title, creator, date(s), extent, and level of description.
The two standards are complementary rather than competing: ISAD(G) supplies an international framework, DACS a U.S. implementation with explicit minimums and explicit deference to institutional judgment. Both are output-neutral.
Separately, the ICA's Records in Contexts offers a contextual, linked model — RiC-FAD, RiC-CM, RiC-O, RiC-AG — that describes records alongside people, places, activities, and events rather than only as a tree. Version 1.0 of the first three parts arrived in late 2023, RiC-O 1.1 in May 2025, and RiC-AG 0.1 in October 2025. It is a real direction of travel; migration and interoperability evidence is still accumulating, so treat it as something to watch and pilot rather than to rebuild around. If you are weighing how these standards stack against transcription encoding and preservation metadata, the layered view of archival metadata standards is the more useful map.
Stage four: writing the notes so a researcher can decide
Two notes carry most of the weight, and conflating them is the most common writing failure.
Administrative and biographical history
This note explains the creator: the life, office, family, or institutional history of whoever is named in the creator element. It is context for why and by whom records exist.
Scope and content
This note explains the records. It answers: what is in this unit; what activities, functions, or subjects does it document; what forms and document types are present; what are the chronological, geographic, and topical boundaries; and what is missing. It should surface concrete names, dates, places, subjects, record forms, languages and scripts, and gaps or exclusions. A title like "Municipal records" answers none of that.
A biographical history is not a substitute for a scope note. Nor is a scope note a restatement of the title. Write it around the decision the researcher is making: is this collection relevant, what does it cover, and what will I actually find when the box arrives?
Betts Coup's 2021 usability study is the closest thing to direct evidence on how these notes get used. Testing three PDF finding aids under task analysis, it found researchers consulted scope-and-content and biographical or historical notes frequently — over 90% identified information in series-level scope notes for two tasks — and relied on collection-level notes more heavily where lower-level description was sparse. It also found that note information gets missed when it sits away from where the user is looking, and that only around 30–40% located and understood access-condition notes (Coup, "The Value of a Note"). The setting was small and specific — PDFs, shared screens, no Ctrl-F — so treat the percentages as design intelligence rather than sector-wide behaviour. The design lesson holds regardless: notes complement lower-level description; they do not replace it. And if your access restrictions are load-bearing, do not bury them in prose.
Stage five: fidelity — what the records say versus what you infer
This is where finding aids quietly go wrong, and where the damage is hardest to detect later.
The professional rule is evidential. Derive every name, date, place, subject, and series label from an explicit feature of the records: an internal index or register, a docket, an endorsement, a heading, or the document text itself. Plausible inference is not evidence. DACS makes the distinction structural by separating a formal title, transcribed according to the applicable transcription standard, from a supplied (devised) title, which the archivist provides where no formal title exists (DACS title element). Labelling a supplied title as supplied is what preserves the boundary between what the records say and what you concluded.
The same discipline applies to normalisation. Silent correction of spelling, dates, place names, or personal names can merge distinct people or distinct villages and conceal what is on the page. Where a normalised form helps retrieval — and it often does — keep the faithful form and add the normalised version as a clearly identified access point or note. Retain the evidence; layer the interpretation on top. The conventions for transcribing old documents cover the mechanics of marking uncertainty and expansion if you need house rules to point staff at.
When you cannot read the source you are describing
There is a stage in this sequence that standards assume and rarely address: someone has to read the records. A series label taken from an endorsement requires reading the endorsement. A date range requires reading dates in the hand they were written in. English wills, French notarial acts, German parish registers, Dutch municipal records, Spanish and Italian administrative papers — all in the Latin alphabet, all requiring different vernacular vocabulary and different paleographic competence. "Latin script" is an alphabet, not a language; the ability to read modern English, or classical Latin for that matter, does not transfer to every vernacular cursive. The National Archives' palaeography tutorial covers English hands of 1500–1800; the Archives nationales runs courses in older French writing of the fourteenth to eighteenth centuries. Both exist because the hand is a genuine barrier, not a formality.
Where that barrier stalls description at scale — a series where the folder headings are illegible, or a bound volume whose internal index is the only route to a container list — handwritten text recognition is the practical way through, provided it does not launder inference into evidence. This is the stage Leo is built for.
Leo's transcription model, ATR-1, reads Latin-script material from roughly the past five centuries — handwritten and printed, in whatever language that alphabet records, with performance strongest in English and strong across French, German, Spanish, Italian, Dutch, Latin and other major European languages. Non-Latin scripts (Greek, Cyrillic, Hebrew, Arabic, Indic, East Asian) are out of scope. It runs as delivered, with no per-collection model training, which matters when the material in front of you is one accession rather than a series large enough to justify building a model for it.
The reason it belongs at this stage specifically is the fidelity requirement above. ATR-1 is trained to transcribe what is on the page rather than to smooth it: strikethroughs, marginal additions, editorial expansions, archaic orthography and u/v interchange survive rather than being silently corrected into modern forms. That is precisely the property a supplied-title decision depends on — you want the endorsement as written, not a plausible modernisation of it.
General chat models fail here in a particular way worth naming. They lean on text prediction with the image downsampled, so their errors arrive as fluent, plausible readings that look like good description and are very hard to catch on review. Transcription errors that are recoverable — a wrong character, checked against the image displayed beside the text — are a different and safer class of problem.
Where a translation into English is needed for internal comprehension, that runs as a separate Transformation writing to its own tab; transcription and translation stay distinct, and the base reading stays untouched. Practically, transcriptions live in a workspace with folder structure, per-document metadata fields (title, creator, date, archive, collection, box, folder, identifier, rights), fuzzy full-text search, and export to TEI XML, Word, PDF, or HTML — output that feeds your descriptive system rather than replacing it. One honest limitation: pages dominated by pre-printed forms with dense handwriting in the fields — ledger and deed-book layouts — are the weakest case, because printed structure can pull attention from the manuscript entries. Budget human review there. If you are weighing this against training your own recogniser, the question of whether you need a custom model is worth settling before you commit ground-truth labour.
Stage six: publishing, and why encoding is not discovery
EAD is an XML standard for encoding hierarchical finding aids, with EAD3 the current schema. It supplies structure and semantics. It does not supply a search interface, and this is the misconception that most often leaves good description undiscoverable.
Publication is its own pipeline: validate or transform your data; expose a readable HTML or PDF representation; index full text, structured fields, or both; maintain links, names, subjects, places, and hierarchy; and test real researcher queries against the result. Structured-field search targets creator, date, subject, and level; full-text search finds the words present in the published text. They are complementary, and neither rescues description that is absent, unindexed, restricted, badly linked, or unharvested.
You do not need to hand-author XML to get there. ArchivesSpace holds hierarchical resource descriptions, exports EAD, and publishes a web finding aid — a benefit the SAA EAD FAQ notes explicitly. AtoM covers hierarchical description, EAD import and export, web presentation and search, and printer-friendly finding-aid output. Contributing to an aggregator such as ArchiveGrid widens cross-repository discovery, but coverage depends on contributor participation, field mapping, data freshness, and working local links — aggregation demonstrates reach, not completeness.
It is worth being blunt about the evidence gap here. There is no defensible current figure for what share of archival description is online, indexed at full text, or represented in an aggregator. OCLC's hidden-collections survey was designed around 275 academic and research libraries in the U.S. and Canada; Harvard's own research guide notes that many finding aids remain undigitised. Those are scope statements, not denominators. If you want to know how discoverable your own holdings are, the only reliable method is to test queries a researcher would plausibly run — which is the discipline behind planning for archive searchability, and it sits alongside the broader digitization and description workflow that a finding aid is only one output of.
What a good finding aid actually does
Every decision in this sequence is a decision about someone else's time. Depth of description trades your hours against a researcher's retrieval requests. Fidelity trades a few minutes of care now against a misattributed name propagating through a dozen published articles. Publishing and indexing decide whether any of it is found at all.
The standards give you the elements and the minimums. What they cannot give you is the judgment about which of your collections deserves file-level attention, which endorsement is evidence and which is a guess, and which researcher question your scope note ought to be answering. That judgment is the craft, and it improves the same way paleography does — by working through collections, recording what you decided and why, and being willing to reopen a description when the retrieval pattern tells you it was wrong.
Frequently Asked Questions
How do you write a finding aid for an archival collection?
Write a finding aid in a fixed sequence. First survey the entire accession to establish what is present, then settle provenance and the best available order, record the order the material arrived in, and triage preservation and access problems. Next choose the level or levels of description the material justifies. Then write the description itself — reference code, title, dates, extent, creator, scope and content, arrangement, access conditions, languages and scripts, rights. Finally publish it in a form researchers can read and machines can index. Arrangement decisions come before description, because re-arranging afterwards means rewriting everything you have written.
What are the required elements of a finding aid under DACS?
DACS single-level minimum requires ten elements: reference code, name and location of the repository, title, date, extent, name of creator (if known), scope and content, conditions governing access, languages and scripts of the material, and rights statements. Administrative or biographical history is an optimum element, not a required one — scope and content carries the weight. A multilevel description applies that minimum at the top level, states the whole–part relationship, and carries appropriate information down the hierarchy, recording creators at lower levels where they differ from the one named above.
What level of description should a finding aid use?
Choose the level as a cost–benefit judgment, not a completeness target. DACS deliberately does not prescribe a correct depth; ISAD(G) supplies the hierarchy of fonds, series, file and item and asks that description move from broad to specific without repeating higher-level information further down. Collection- or series-level description is usually adequate for large, homogeneous or low-risk material. File- or item-level work is justified where titles do not communicate content, where research or evidential value is high, where rights and access decisions must be made unit by unit, or where access benefit exceeds processing and ongoing maintenance cost.
What is the difference between a scope and content note and a biographical history?
A biographical or administrative history explains the creator — the life, office, family or institutional history of whoever is named in the creator element. A scope and content note explains the records: what is in the unit, what activities, functions or subjects it documents, what forms and document types are present, what the chronological, geographic and topical boundaries are, and what is missing. One is not a substitute for the other, and neither is a restatement of the title. Write the scope note around the decision a researcher is making: is this relevant, and what will I find when the box arrives?
What if you cannot read the handwriting in the records you are describing?
Description depends on reading the records — a series label taken from an endorsement requires reading that endorsement, and a date range requires reading dates in the hand they were written in. Where the hand stalls description at scale, handwritten text recognition is the practical route through. Leo's model, ATR-1, reads Latin-script material from roughly the past five centuries, handwritten and printed, and is trained to transcribe what is on the page rather than smooth it — strikethroughs, marginal additions, archaic orthography and u/v interchange survive. That matters here, because a supplied-title decision depends on the endorsement as written, not a plausible modernisation.