Passenger List Transcription: Reading Ship Manifests Column by Column
Passenger-list transcription of 1907 U.S. ship manifests column by column across both pages, joining them by line number so every field stays attached to the correct passenger.
Leo Team
August 10, 2026

Contents
Passenger list transcription is the work of copying a ship manifest row by row and column by column, so that every value stays attached to the passenger it describes. This guide walks through the 1907 arrival form field by field, explains why the second page is where most researchers lose the thread, and sets out a working method you can still defend years later.
On a manifest, meaning is relational: an age, a residence, a sum of money, or a named relative is only useful if you can prove which passenger it describes. That is why transcribing a manifest is a different job from transcribing a will or a letter — the hard part is not only reading the handwriting, but preserving the grid. Work from the original image, transcribe the full width of the row across both pages, and treat the online index as a pointer to that image rather than as the record itself.
What a manifest actually is, and what survives
A passenger manifest is a carrier and immigration document organized as a list. Each passenger occupies a logical row; ruled columns assign separate facts to that row. It was compiled at the point of departure, not at the port of arrival — a detail that matters, and one we will return to.
The federal record begins at a fixed date. The National Archives records that the U.S. government did not require vessel masters to present a passenger list until January 1, 1820, and holds foreign-arrival records from approximately 1820 through December 1982, with gaps, arranged by port. There are two exceptions to the 1820 rule: New Orleans arrivals 1813–1819 and Philadelphia arrivals 1800–1819. Records older than 75 years are publicly available; more recent ones are restricted and handled under privacy and FOIA rules.
Genealogical writing commonly draws a line between pre-1891 customs lists and post-1891 immigration lists, and you will see confident chronologies of every column added in 1893, 1903, and 1906. Treat those chronologies with some caution unless you can check them against archival form images; the widely circulated 22-column description of the 1893–1906 single-page format comes from a secondary extraction aid rather than a primary archival source. What is firmly documented is the 1907 form, and it is the one worth learning in detail.
The 1907 form, column by column
The National Archives publishes a specimen of the List or Manifest of Alien Passengers for the United States, required under the Act approved February 20, 1907 and labeled "February 1907 to February 1917." It is the clearest single reference for what the columns say and, just as importantly, in what order they appear.
Page 1: identification
The left-hand page opens with the administrative and identifying fields:
- No. on List — the line number. This is the key that binds page 1 to page 2. Record it always.
- Full Name — split into Family Name and Given Name.
- Age — in years and months.
- Sex; Married or Single; Calling or Occupation.
- Able to Read / Write — two separate sub-columns.
- Nationality (Country of which citizen or subject).
- Race or People — a period administrative classification, not a modern statement of ethnicity. Transcribe the term the clerk wrote; interpret it separately.
Then the fields that begin to earn their keep for research:
- Last Permanent Residence (country, city or town) — the place the passenger reported as their recent permanent home. It is not necessarily a birthplace, and conflating the two is one of the commonest errors in family trees.
- The name and complete address of nearest relative or friend in country whence the alien came — often the single most valuable line on the sheet for pushing a line back across the Atlantic, because it names a person and a street or village. Read the header literally: it says "relative or friend." It is a lead to a person and a place, not proof of kinship.
- Final Destination (state, city or town) — where the passenger intended to go, which is not evidence of where they eventually settled.
Page 2: the continuation sheet
From 1907 the standard one-page manifest expands to two pages, as the Statue of Liberty–Ellis Island Foundation's search guide explains, with additional questions on the continuation sheet. The continuation carries:
- Whether having a ticket to such final destination.
- By Whom was the Passage Paid? — with sub-fields distinguishing the alien's own payment, payment by another person, and payment by a corporation, society, municipality, or government. This distinction is a research lead in itself: a passage paid by a society or a corporation implies a sponsorship or labor context worth chasing.
- Whether in possession of $50, and if less how much?
- Whether ever before in the United States; and if so, when and where? — a direct pointer to an earlier crossing, and one of the most frequently missed fields in family research.
- Whether going to join a relative or friend, with name and complete address.
- Institutional and exclusion questions: prison or almshouse, care for the insane, support by charity; whether a polygamist; whether an anarchist; whether coming under an offer, solicitation, promise, or agreement to labor.
- Condition of health, mental and physical; deformed or crippled, with nature, duration, and cause.
- Height (feet and inches), Complexion, Color of Hair, Eyes, and Marks of identification.
- Place of birth (country, city or town).
- Passport, visa, or re-entry permit information where applicable, and prior deportation.
Every one of these is a contemporaneous declaration, recorded by a clerk, from a traveler who may not have shared the clerk's language. They are evidence of what was stated, not settled fact. Transcribe them as written and reason about them afterwards.
Why the width of the row is the whole problem
Here is the structural trap. Page 2 is a continuation of page 1 and does not list passenger names. The Foundation's guide adds a practical warning that catches almost every researcher once: page 2 often appears first in the scroll of digitized frames, and page 1 and page 2 will always be adjacent — if page 1 is frame 333, page 2 is frame 332 or 334.
So the sequence "find the name, read the fields next to it, stop" quietly discards half the record. Worse, the sequence "find an interesting line on the right-hand page and attribute it to your ancestor" produces a confident error: the money carried, the relative joined, and the physical description assigned to the wrong person. You join the two pages by line number and manifest sequence, never by eye and never by vertical position alone.
This is not only a genealogist's discipline problem. It is the same problem that makes manifests genuinely hard for machines. Document layout analysis has to segment a page into logical objects — lines, words, cells, rules, background — before recognition happens at all, and research on historical registers is explicit that recognition of the table structure, columns and headers, is the prerequisite for row detection and handwritten text recognition. A comparative study of Swedish handwritten tabular records found existing OCR tooling simply could not perform the layout analysis, with skew, curl, and broken or discontinuous rules as material obstacles. On microfilmed manifests you have all of it: warped rules, faded ink, printed grid over cursive entries, and a row geometry that shifts across the sheet.
The practical consequence: a recognizer can return perfectly plausible characters and put them in the wrong cell. That is a worse outcome than a visibly garbled character in the right cell, because nothing about it looks wrong.
The index is a finding aid, not the record
Most people meet a manifest through a search box. The Ellis Island / Port of New York ecosystem is large — the Foundation describes a searchable database of 65 million Port of New York arrival records covering manifests from 1820 to 1957, a collection figure rather than a count of individuals — and FamilySearch and Ancestry index a great deal besides.
Use them. Then open the image. Three reasons.
Indexes carry acknowledged transcription risk. The Foundation states plainly that there is always room for spelling error in the handwriting or type on a manifest, or in the digital transcription. No reliable published error rate exists for the major passenger indexes; what exists is a documented correction mechanism and an acknowledged risk. FamilySearch permits corrections on many image-linked indexes, though not every indexed image is editable and you cannot edit another user's correction.
Field scope varies by collection. A name/age/ship hit does not mean the whole wide row has been transcribed. The genealogically valuable fields sit physically on the continuation page, and there is no guarantee a given index reached them.
Annotations are not indexed at all. Marginal and stamped material is often the most consequential thing on the sheet. NARA documents manifest annotations including "N.O.B." (Not on Board), "did not sail," and "not shipped" in the left margin, and notes that officials would cross off duplicate entries to leave one official record per person. The Foundation's guide covers hospital, discharged/deported/died, NAT, re-entry-permit, and special-inquiry markings. A person can appear on a manifest and never have crossed. Read the margins.
This is the same discipline that governs FamilySearch full-text search as a discovery layer rather than a transcript, and it belongs in any broader approach to transcribing genealogy and family history records.
The name myth, and what to search instead
The persistent story that officials at Ellis Island changed immigrants' names does not survive contact with the record. The Foundation states that, for sea and air arrivals, arrival records were filled out at the point of departure; clerical errors were possible, but immigrants were not given new names on arrival. What that means for transcription is concrete: the name on the manifest reflects a European clerk, a European spelling convention, and often a language other than English. Search the ethnic form as well as the later Anglicized one — Giuseppe as well as Joseph, Zsuzsanna as well as Susan.
When you transcribe, copy the spelling on the page exactly, however wrong it looks. Your reconciliation of "Kowalczyk" to "Kovalchik" belongs in a note, not in the transcription. The same principle runs through all faithful genealogy record transcription: transcribe first, interpret second, and keep the two visibly separate.
Cross-check with the departure series
The arrival manifest is one side of a crossing. Several departure and counterpart series exist, each with its own layout, language, and conventions:
- Hamburg passenger lists, covering millions of Europeans who left via Hamburg between 1850 and 1934, split into direct and indirect lists, each with its own handwritten index.
- UK Board of Trade BT 27, outward lists for passengers boarding at UK and Irish ports 1890–1960.
- Canadian Form 30A, recording people arriving in Canada by sea 1919–1924 — and note that Library and Archives Canada says those records are not name-searchable and must be browsed on microfilm.
These are independent evidence, not interchangeable copies. Where a departure list and an arrival manifest disagree on an age or a village, you have a discrepancy to investigate — not automatic proof of a name change. And where the departure list is in German, remember that transcription and translation are two separate operations: you want a faithful source-language text first, and an English rendering as a clearly labeled second artifact. Collapsing them, as many searches for an "old handwriting translator" do, loses the citable original.
A working method for one manifest
- Capture or download both pages at full resolution. If you are photographing microfilm or a bound volume yourself, get the geometry right — the field method for photographing archival documents applies to reading-room microfilm readers too. Skew and curl degrade every downstream step.
- Locate the line number first. Not the name — the number. It is your join key.
- Confirm the frame order. Establish which frame is page 1 and which is page 2 before you read a single value from the right-hand sheet.
- Transcribe the header row. Copy the printed column headings verbatim from the form before you copy any entries. Later you will thank yourself for knowing whether column 14 asked about a ticket or about $50.
- Transcribe your ancestor's row across both sheets, cell by cell, in column order, marking any cell you cannot read rather than guessing. Blank means blank; illegible means illegible; the two are different findings.
- Transcribe the neighbouring rows. Families and village cohorts travel together and are usually consecutive. This is the step most researchers skip and most often regret.
- Read and record the margins and stamps against the line number.
- Verify the decisive columns against the image a second time — names, ages, dates, place names, and the relative's address. These are the fields where an error propagates furthest.
Where machine transcription helps, and where it needs watching
For a single ancestor's row, careful manual transcription is perfectly adequate. The case for machine assistance arrives when you are working through a whole sheet, a family group scattered across several sailings, or a run of manifests for a village chain migration — the point where reading 30 columns by hand for 30 lines stops being research and starts being data entry.
Two cautions are specific to this material. First, do not paste a manifest into a general chatbot and accept the answer. On a dense grid, a general-purpose model will produce fluent, well-formatted output that reads as authoritative and can silently reassign values between columns or invent a plausible occupation — the fluent-but-wrong failure mode that is far harder to catch than a garbled character. Second, be honest about what any tool can and cannot do with pre-printed forms.
That second caution applies to Leo as much as anything else, and it is worth stating plainly. Leo's transcription model, ATR-1, is built for Latin-script material — any language written in the Latin alphabet, which covers English manifests, German Hamburg lists, and Italian or Dutch departure records alike — and it transcribes what is on the page rather than smoothing archaic spelling and clerk's abbreviations into modern prose. It reads without any per-corpus model training, which matters when your material is a handful of sheets rather than a corpus large enough to justify building a model. It preserves tables in output, and the original image sits beside the transcription in the editor, which is the layout you need for cell-by-cell verification. But heavily pre-printed forms with dense handwriting in the cells are Leo's known weak spot: the model can favor the printed headers over the manuscript entries, and complex tabular layouts vary in quality. On a manifest, use it to get a first-pass reading of the page quickly, then verify the decisive columns yourself against the image — which is what the method for verifying transcription accuracy asks of any machine output, and what you should demand of any tool that claims to read a grid.
The realistic standard for automated table reading is worth internalizing here. The historical-register study cited earlier reported a mean cell match of 88.28% on its own corpus of death registers — a genuine result, on 70 images, from a different record series, and not a passenger-manifest benchmark. No defensible accuracy claim exists for any system on Ellis-era manifests without a labeled test set. Treat structural preservation, not raw character accuracy, as the measure that matters: a page where every cell is in the right place with three uncertain characters is more usable than a page of clean text in the wrong columns.
What you are actually building
A transcribed manifest row is not a lookup result. It is a small, dated, sourced statement of what one clerk in one European port wrote down about one traveler on one day — with a line number that ties it to the sheet, a page reference that ties it to the film, and your own marks where the ink defeated you. Kept that way, it holds up when a cousin's tree disagrees with it, when a later record contradicts an age, or when you come back in five years having forgotten why you were sure.
The columns were designed to be read across. Read them that way, keep what you cannot resolve visibly unresolved, and the row will still be telling you the truth long after the search index has been re-indexed and the frame numbers have changed.
Frequently Asked Questions
What is passenger list transcription?
Passenger list transcription is the work of copying a ship manifest row by row and column by column, so every value stays attached to the passenger it describes. It differs from transcribing a will or a letter because meaning on a manifest is relational: an age, a residence, a sum of money or a named relative only helps if you can prove whose row it sits in. Work from the original image, transcribe the full width of the row across both pages, copy the printed column headings verbatim, and treat any online index as a pointer to the image rather than the record itself.
How far back do U.S. passenger arrival records go?
The federal requirement begins on 1 January 1820, when the U.S. government first required vessel masters to present a passenger list. The National Archives holds foreign-arrival records from roughly 1820 through December 1982, with gaps, arranged by port. Two exceptions predate 1820: New Orleans arrivals for 1813–1819 and Philadelphia arrivals for 1800–1819. Records older than 75 years are publicly available; more recent ones are restricted and handled under privacy and FOIA rules. For crossings before 1820, look to departure and counterpart series rather than U.S. arrival manifests.
Why does page 2 of a ship manifest matter, and how do I find it?
From 1907 the manifest expands to two pages, and page 2 carries the fields researchers most want: who paid the passage, whether the passenger held $50, whether they had been in the United States before, who they were going to join, health and physical description, and place of birth. Page 2 does not repeat passenger names, so you join it to page 1 by line number, never by eye or vertical position. In digitized film, page 2 often appears first in the scroll; the two frames are always adjacent, so if page 1 is frame 333, page 2 is frame 332 or 334.
Did Ellis Island officials change immigrants' names?
No. Arrival records for sea and air arrivals were filled out at the point of departure, not on arrival, so officials at Ellis Island were checking names against a list rather than assigning new ones. Clerical errors were certainly possible, but the name on the manifest reflects a European clerk, European spelling conventions, and often a language other than English. For searching, that means trying the ethnic form as well as the later Anglicized one — Giuseppe as well as Joseph, Zsuzsanna as well as Susan — and transcribing the spelling on the page exactly, with your reconciliation kept in a note.
Can AI or OCR transcribe ship manifests accurately?
Not reliably enough to accept unverified. Ship manifests are among the harder documents for machines: table structure has to be recognised before rows and handwriting can be read, and microfilmed manifests bring skew, curl, broken rules, faded ink and a printed grid over cursive entries. A recognizer can return plausible characters in the wrong cell, which is worse than a visibly garbled character in the right one. One study of historical death registers reported a mean cell match of 88.28% on its own small corpus, but no defensible accuracy figure exists for any system on Ellis-era manifests. Use machine output as a first pass, then verify the decisive columns against the image.