“A Probabilistic Finding Aid”: How Professor Thiago Krause is Using Leo for a Global History of Salvador da Bahia

Thiago Krause, a historian of Brazil and the Atlantic world at Wayne State University, is writing a global history of Salvador da Bahia drawn from more than a hundred archives—a scale made possible by handwritten text recognition and AI agents, which have sifted hundreds of thousands of manuscript images for him. But the payoff has not been less reading. He now spends more time with manuscripts than he ever did. We spoke with him about how tools like Leo are changing the way he works and about what he refuses to outsource to machines.

Jon Cooper

Founder and CEO · July 31, 2026

Thiago Krause, historian of Brazil and the Atlantic world at Wayne State University
Contents

Thiago Krause is a historian of Brazil and the Atlantic world at Wayne State University, where he writes about slavery, inequality, and commerce between 1450 and 1850. Born and raised in Rio de Janeiro, he earned his Ph.D. from the Universidade Federal do Rio de Janeiro in 2015 and taught at UERJ and Unirio before moving to Detroit. He is currently writing a global history of Salvador da Bahia, reconstructed from more than a hundred archives across Europe, the Americas, and beyond — a project that has made him an early and unusually sophisticated adopter of handwritten text recognition and AI-assisted archival research.

João Teixeira Albernaz, Planta da Restituição da Bahia, 1631: a bird's-eye view of Salvador and the fleet that retook it from the Dutch in 1625

João Teixeira Albernaz, “Planta da Restituição da Bahia,” 1631, Wikimedia Commons.

Salvador da Bahia as a global port

To set the stage for our readers, could you introduce yourself and tell us about your research? What drew you to this specific topic?

I am a historian of colonial Brazil and the early modern Atlantic world. My current book project is a global history of Salvador da Bahia from the sixteenth century to 1763, when it ceased to be the capital of Portuguese America. I am interested in Salvador not only as a Brazilian city, but also as one of the central hubs of the early modern world: a port with direct and indirect connections to American plantations, African slave markets, European consumers, Asian textiles, Indigenous North America, the Mediterranean, the Río de la Plata, and the Indian Ocean.

The book does two things at once. First, it restores Salvador to the place it occupied in its own time. For much of the seventeenth and eighteenth centuries, it was one of the largest cities in the Atlantic seaboard of the Americas and its most important slaving port. Second, it uses Salvador to rethink larger histories of globalization, capitalism, consumption, and slavery. I argue that early modern globalization was not simply directed from London, Amsterdam, Paris, or Lisbon; it was equally shaped by colonial merchants, African rulers and consumers, Indigenous traders, enslaved workers, imperial officials, and many others whose choices and coercions structured commercial life.

What drew me to the topic was, in part, a mismatch between historical importance and historiographical visibility. Salvador appears everywhere once one starts looking across archives: in French merchant correspondence, English commercial papers, Dutch company records from West Africa, Spanish American criminal investigations, Genoese notarial deeds, Portuguese imperial documentation, and even sources about Macau and China, among many others. Yet colonial Salvador is still much less visible than it should be, especially outside Luso-Brazilian historiography. I became increasingly interested in using one place to write an empirical global history, rather than writing a largely synthetic overview in which places disappear into abstract networks.

A recurring example is Brazilian tobacco. It was cultivated and processed in Bahia by enslaved workers, but became desirable in West Africa, Spain, Italy, Central Europe, China, and among Indigenous consumers in North America. Its history shows how a commodity could travel globally not only because Europeans imposed it, but because consumers in many places developed specific preferences and refused substitutes. At the same time, those preferences were tied to coercion: tobacco helped purchase enslaved Africans, who in turn grew that same crop. That tension, between choice and violence, between local agency and structural domination, is at the center of my work.

The workflow before AI

When you presented your job talk a few years ago, your research already drew from 50 archives. Now that number has expanded to well over 100 scattered repositories. Could you describe what a typical historical research workflow looked like before you integrated tools powered by recent developments in AI? What were the main bottlenecks when dealing with thousands of pages of handwritten archival records?

Before these tools, the workflow was much more linear and much more constrained by human reading time. I would identify a promising archive, series, or volume; travel there or use digitized material; photograph or download as much as possible (or, until a few years ago, transcribe and take notes in situ); organize the images manually; and then inspect the material page by page, taking notes in Word. If a volume looked promising, I might spend days or weeks reading it. If it looked only vaguely relevant, I often had to make a judgment call: invest the time and perhaps find nothing, or move on and risk missing something important.

That is the basic problem of scattered archival research. To trace transimperial connections, one cannot rely on documents being indexed under Bahia or Brazil. They appear in correspondence about the French Caribbean, Dutch West Africa, English tobacco policy, Spanish American contraband, Genoese commerce, or the North American fur trade. Catalogues rarely capture that level of detail. A document can be catalogued as a routine diplomatic letter and still contain a crucial paragraph on Bahian sugar, tobacco, gold, or slaving voyages. So the bottleneck was discovery.

Handwriting compounded the problem. Early modern corpora are multilingual, inconsistent, and messy. Spelling is unstable; the same place or person may appear under several forms; and a single “item” can contain a letter, enclosures, summaries, copies of earlier correspondence, and later docketing. Before HTR and recent LLM tools, the only reliable method was slow reading by someone with enough paleographic and historical knowledge to know what mattered.

The result was a permanent tension between breadth and depth. I could read closely, or I could cover more ground, but doing both across multiple countries and dozens of archives was extremely difficult. Digitization had already changed the scale of what was accessible, but it created a new problem: there were suddenly far more images than any single historian could read. LLMs do not solve the problem of interpretation, but they have changed the scale at which initial discovery and triage can happen.

From bespoke models to out-of-the-box HTR

You mentioned that you spent two years working primarily with Transkribus, training bespoke machine-learning models for difficult handwriting; the process required upfront work on layout recognition but paid off for highly complex scripts. More recently, though, you adopted Leo for the bulk of your Latin-script documents. Could you talk about what prompted that shift? From a practical standpoint, what are the distinct advantages of moving away from training custom models toward a platform that reads Latin-script documents immediately out of the box?

Transkribus was essential for me because some of my sources are genuinely difficult. When the handwriting is especially hard, bespoke models can still be extremely useful. The most important examples for my research are Dutch and Brazilian notarial records, along with one particular set of merchant correspondence. LLMs still cannot read those reliably. For difficult scripts, training a model can be worth the investment, particularly if it is a collective enterprise dedicated to a large collection.

But the investment is real. You need to segment pages, correct transcriptions, train, test, retrain, and then repeat the process for a different hand or archive. That makes sense when the corpus is bounded and difficult enough to justify the cost. It makes less sense when you are dealing with tens of thousands of documents in many languages and from innumerable hands, where the main problem is not achieving a perfect transcription of one corpus but making a vast amount of material searchable quickly enough to know what deserves close attention.

That is what prompted the shift. For the bulk of my materials, Leo lowered the threshold between image and usable text. Instead of asking, “Can I train a model for this archive?” I can ask, “Can I process this corpus now, preserve the images and metadata, and make it searchable enough for discovery?” This is a different research regime, one where perfection matters less than discoverability. After all, I still have to read and check the images for every document I am going to use, to be sure of what they actually say.

The practical advantage is speed: not the speed of skipping the reading, but the speed of getting from a box, a volume, or a folder to a searchable research environment. That matters enormously for scattered research. I do not need every line to be perfect in order to identify a memorandum about Bahian tobacco in Genoa, a French letter about smuggling in Salvador, or an English comment on Brazilian sugar. I need a reliable enough first pass that points me toward documents worth human verification. The result is that I work with a much larger source base, and in the end spend more time reading manuscripts than I ever did.

Again, I would not frame this as replacing custom models entirely. For some material, especially difficult hands or non-standard layouts, bespoke training remains important. But for large bodies of Latin-script material, an out-of-the-box system changes the economics of research. Transkribus has its Titan II model, which is quite good, but Leo usually produces more readable output and does not require separate layout analysis, so it is generally more efficient. It allows the historian to reserve the most intensive labor for interpretation, verification, and close reading, rather than spending months merely creating the conditions for search, as I did for the first seventeen years or so of my career.

What a single historian can now do

Many scholars remain hesitant to build large digital corpora because they assume the technical or computational hurdles are still too high. In your experience, how does having access to a platform that mostly works out of the box change what a single project can achieve? What shift in thinking is required for historians to take advantage of this?

The main change is that large corpora become feasible for individual scholars, not only for funded teams. That does not mean every historian should build an enormous corpus, or that scale is automatically good. But it does mean that the old assumption, that a single scholar can only work with what they can personally read line by line from the start, is no longer quite right.

For my project, this has meant that I can work at two levels simultaneously. I can build a broad corpus across many archives, languages, and imperial contexts, and then use tools to identify the smaller subset of documents that require traditional historical attention. The machine helps with reconnaissance. The historian still has to decide what the corpus is, what counts as relevance, how to handle uncertainty, what must be checked against the image, and how a document fits into a larger argument.

The required shift is to stop thinking of HTR as either a perfect transcription or a failure. For many research purposes, it is better understood as a probabilistic finding aid. It lets us ask better questions of uncatalogued or poorly catalogued material, but it does not remove the need for archival judgment: machine output should function as a map of possible evidence, not as evidence itself.

There is also a methodological shift. Historians need to think more explicitly in terms of bounded corpora, provenance, metadata, selection criteria, and verification. What exactly did I process, and what did I exclude? Which languages or hands are underperforming, what kinds of errors is the model likely to make, and what would count as a false negative? These questions are not alien to historical research; they are versions of source criticism. But digital corpora force us to make them more explicit.

In that sense, the most important change is intellectual, not simply computational. These tools make it possible to write histories that are less constrained by national archives and imperial catalogues. But they only do so when historians design the corpus carefully and resist the temptation to treat search results as transparent access to the past.

Querying uncatalogued material with AI agents

You recently used Leo to transcribe around seven thousand merchant letters, but this is more than you could possibly have time to read individually. Could you walk us through how you combine handwritten text recognition and AI agents to query uncatalogued material? How does this capability change the possibilities of archival discovery, particularly for scattered or poorly indexed collections? And what are the risks to be aware of in designing these processes?

The workflow begins before transcription. First, I try to define a bounded corpus: a particular archive, series, volume, merchant collection, date range, or set of images. I want to preserve original order, image numbers, and archival metadata, because without provenance the results become much less useful. Then I transcribe the material with Leo, export the output, and use LLM agents to index it in a spreadsheet.

The agents are not asked to “write the history.” They are asked to help identify candidates. For example, I might ask them to highlight letters concerning Brazilian tobacco, sugar prices, West African trade, Portuguese merchants, contraband, Bahia, Pernambuco, Lisbon reexports, or specific people and places, always with variant spellings and multilingual terminology in mind. A useful output is a table of candidate letters, not a synthetic essay: dates, image references, snippets, confidence, and a short explanation of why the item may matter.

Likewise, I’ve recently begun using agents to work through hundreds of thousands of images in multiple collections, identifying possible candidates for transcription and examination. I’ve lost count, but I’m sure they have processed more than 300,000 images in three months. The collections include Portuguese Inquisition denunciations and imperial records, French diplomatic and imperial correspondence, Antwerp notary books, Dutch company records, and more. I’m doing that because the cost of transcribing all those images with Leo would be astronomical, and visual triage costs only roughly 1% as much; I will nevertheless use Leo to transcribe the tens of thousands of selected images.

Then comes the essential step: I return to the document. I check the image, the HTR, the surrounding pages, and the document boundaries. Is this really one letter, or is there an enclosure? Is the key phrase in the body of the letter or only in later docketing, and did the model misread a place name? Does the passage actually concern Brazil or is it a false positive? Only after that do I decide whether the document enters my research notes, databases, or selected corpus for close reading.

This changes archival discovery because uncatalogued material becomes queryable at a conceptual level. Merchant letters are a good example. A catalogue may tell you the name of a merchant house and the dates of its correspondence, but not that three letters in the middle of a volume discuss Brazilian tobacco in Naples, rumors about Bahian sugar in Hamburg, or attempts to bypass Portuguese restrictions in Bahia. HTR plus agents make it possible to find those needles without pretending that the entire haystack has become transparent.

A 1719 merchant letter in French from Lisbon to Bristol, mentioning a possible slaving voyage from Madagascar to Brazil

Arquivo Nacional da Torre do Tombo, Lisbon, Junta da Administração do Tabaco, Livro 185: 153, Willem de Bruijn & Paulo Cloots, Lisbon, to Christopher Schutter, John Corsley and others, Bristol, January 7, 1719. In this letter written in French, two Dutch merchants mention the possibility of a slaving voyage from Madagascar to Brazil.

The risks are serious. HTR can miss the most important word on a page. LLM agents can and do overinterpret, normalize away uncertainty, or produce plausible but wrong summaries. They can also reinforce the researcher’s expectations: if I only ask for Brazil, I may miss documents that matter precisely because they describe the same circuit from an African, Spanish, Dutch, or French perspective. And absence is especially dangerous. The fact that an agent did not find something does not mean it is not there.

So the workflow has to be designed defensively. I try to keep the corpus bounded, preserve provenance, inspect results against images, sample negative results when possible, and avoid using OCR silence as evidence of absence. The point is not to automate trust but to create a better system for deciding where human attention should go.

What historians must not outsource to machines

How do you decide what to outsource to an automated system and what must stay as the province of the human scholar? If a specialized model can transcribe the manuscript and agents can find relevant passages in the document, what, in other words, is the core of the historian’s craft that we must refuse to outsource to machines?

I outsource tasks that are repetitive, mechanical, or preliminary: downloading, file organization, first-pass transcription, rough search, candidate identification, table creation, and sometimes the extraction of names, dates, and places. These are labor-intensive tasks, and historians have always used tools for them: calendars, indexes, catalogues, research assistants, databases. AI changes the scale and speed, but not the basic distinction between assistance and judgment.

What I do not outsource is the decision about what a document means. A model can identify a passage about tobacco, but deciding why that passage matters for the history of Atlantic slavery, consumer preference, or Portuguese imperial governance takes the historian’s framework. The same is true of a reference to Bahia: nothing in the model reliably distinguishes a passing mention from a document that matters analytically. And a model can summarize a letter without grasping the archive’s silences, the institutional context of the collection, or the historiographical stakes of the evidence. Despite improvements in internal knowledge and context windows, these systems remain far from being able to hold all the context historical research requires.

The core of the historian’s craft is judgment under conditions of incomplete and biased evidence. Historians decide which questions are worth asking, which archives might answer them, what counts as evidence, how much weight to give a source, and how to build an argument without flattening contradiction or uncertainty. That requires knowledge of language, paleography, institutions, historiography, and human behavior. It also requires ethical judgment, especially in histories of slavery, where the archive was often produced by enslavers, merchants, judges, and imperial officials.

There is also something more basic: historians must preserve the capacity to be surprised by the archive. LLMs are trained to identify patterns and tend to average things out, so they often reproduce what is already known. Automated systems work best when we define what we are looking for. But some of the most important discoveries come when a document does not fit the categories we brought to it. If we only use machines to retrieve expected answers, we risk turning archives into confirmation devices. The scholar’s role is to keep the research open to anomaly, context, and contradiction.

So I do not see the line as “machines transcribe, humans write.” It is more fundamental than that: machines can help us move through evidence, but they cannot assume responsibility for historical understanding.

Teaching students to avoid skill atrophy

You’ve argued elsewhere that these tools augment the work of those who already possess historical judgment, but risk producing atrophy in students who are still building those skills. How can we teach historical research in a way that makes use of these new technologies while preserving the learning process?

I think the key is sequencing. Students can use these tools because they are becoming part of historical research, though many excellent historians choose not to and will continue not to; that is no judgment on their work. But students should not encounter them first as shortcuts. They need to learn what the tools are shortcutting.

That means students should still do basic archival work manually: read manuscript images, struggle with handwriting, identify document boundaries, compare catalogue descriptions with actual documents, transcribe a page, build a small source table, and explain why a document is or is not relevant. Only then does it make sense to introduce HTR or LLM agents and ask: what did the tool get right, what did it miss, and why?

One useful assignment would be comparative. Give students a small manuscript corpus. First, have them inspect it manually and build a basic index. Then have them run HTR or query LLM-generated transcriptions. Finally, have them compare the results: false positives, false negatives, mistranscriptions, missed marginalia, misunderstood names, and misleading summaries. The point is not to make them distrust the tools but to help them understand the tools as fallible instruments that require historical supervision.

We should also teach prompt design as a form of methodological design. A prompt is not just a technical instruction; it encodes assumptions about relevance, evidence, categories, and exclusion. If students ask a bad question, the machine may give them a tidy but useless answer. So they need to learn how to define a bounded corpus, specify inclusion and exclusion criteria, preserve provenance, and require citations back to images or documents.

The danger is that students may confuse fluency with understanding. A generated summary can sound authoritative even when it rests on a misread word or a shallow grasp of context. To prevent that, I would require students to cite original documents, not LLM outputs; to verify claims against images; and to explain uncertainty rather than hide it.

In the end, these tools should be taught as part of archival literacy, not as a replacement for it. The goal is not to train students to avoid the hard parts of historical research but to help them reach more sources while becoming more explicit, disciplined, and critical about how they use them.

Share this article

© 2026 Leo Technologies Limited. All rights reserved