Leo Collections: Our Plan to Build a Universal Archive

The world's documentary heritage lies scattered across thousands of local archives, paywalled databases, and hard drives. Most of it has never been catalogued, still less digitized. Leo Collections is an ambitious attempt to draw it together. We're building a crowdsourced, machine-readable archive of the world's textual past, in four steps: gathering images, transcribing them, making collections speak to one another, and surfacing the patterns held within.

Jon Cooper

Founder and CEO · July 5, 2026

Detail from the Catalan Atlas, c. 1375, showing the Sultan of Delhi (top) and the king of Vijayanagar (bottom). Compiled in Majorca by Cresques Abraham from merchants' reports, travellers' tales and older maps, it gathered a world never before seen in a single view
Contents

The world's documentary heritage lies scattered across thousands of local archives, paywalled databases, and hard drives. Most of it has never been catalogued, still less digitized. Leo Collections is an ambitious attempt to draw it together. We're building a crowdsourced, machine-readable archive of the world's textual past, in four steps: gathering images, transcribing them, making collections speak to one another, and surfacing the patterns held within.

Patrick Manning, Mellon Professor of World History at Pittsburgh and a former president of the American Historical Association, spent much of his career on a problem his colleagues deemed unrealistic, even mad: to develop the infrastructure needed to write a comparative history of the world. Such a history, he argued, could not be built on traditional scholarly habits. It needed shared databases, standardized metadata, and systems interoperable enough to hold evidence from different languages, societies, and documentary traditions within one view. Through the World History Network and the Collaborative for Historical Information and Analysis at Pittsburgh, he brought historians and data scientists together to pool evidence about population, migration, and social structure.

Though it inspired much fruitful work, Manning never quite realized his dream. The world's archival heritage remains scattered across miles of shelving in national and local repositories, unmarked boxes in municipal offices, parish registers, ship's logs, and census records strewn across jurisdictions in specialist collections. The fraction that has been catalogued is described under wildly inconsistent schemes, and the even smaller portion that has been digitized is dispersed across hard drives, brittle links, and project sites that vanish when funding expires.

Every dimension of the task cuts against unification: the sheer volume across thousands of institutions; the heterogeneity of sources that resist clean transcription and tidy standardization; metadata schemas devised in different cultural and computational worlds; and a professional reward structure that prizes novel findings over the unglamorous work of building shared infrastructure.

Despite the practical difficulties, the potential rewards of a global archive are immense. You can catch a glimpse of them where a historian has linked far-flung records by hand. Natalie Zemon Davis, in Trickster Travels, reconstructed the life of al-Hasan al-Wazzan, a Muslim diplomat from Fez, seized by Christian pirates and presented to a pope, who became one of Renaissance Europe's principal informants on Africa and was baptized as Leo Africanus. Tracing those connections across languages and continents took Davis years. In a truly interconnected archive, they could surface in seconds.

Such a world archive would change what kinds of questions could be asked about the past. Patterns now buried across dispersed ledgers and unindexed manuscripts would come into view, from the webs of migration and exchange that tied distant societies together to the movement of ideas across linguistic and imperial borders. Fernand Braudel's attentiveness to the longue durée, the deep structures of geography and demography beneath the surface of events, was constrained by the limited evidence available to test it. Where Braudel and the annalistes were forced to work from selective samples and informed conjecture, a universal archive would offer a more secure evidentiary foundation for the same panoramic view.

Ramelli's bookwheel, from Le diverse et artificiose machine: an imaginary machine that held a dozen open books at once and turned them sequentially past the reader's gaze

Ramelli's bookwheel, from Le diverse et artificiose machine: an imaginary machine that held a dozen open books at once and turned them sequentially past the reader's gaze. Source ↗.

A potted history of the universal archive

Manning did not invent the idea of gathering knowledge into a single ordered store. The practice has no clear beginning at all, since just about every literate culture has cared to preserve its written heritage. Some fifteen thousand clay tablets survive from the royal palace at Mari, on the Euphrates, dating to the second millennium BC. Such gathering tends, soon enough, to lead into dreaming of gathering everything, of a one, unified emporium containing all possible knowledge. Thomas Richards evocatively describes the "imperial archive" as a kind of utopian fantasy, of knowledge unified in the service of state power.

That vision took firmer institutional shape after the turn of the first millennium, when written documents multiplied across Europe. Charters that defined legal relationships were produced in growing numbers, rulers gathered information about their subjects as a means of governing them, and a class of trained scribes, secretaries, and copyists emerged to produce and keep the new profusion of records. In Markus Friedrich's retelling, the cartulary, an organized book collecting a monastery's or a noble house's important documents in one place, and the register, a systematic copy of documents issued or received, were critical early steps toward the deliberate, organized preservation of records.

By the end of the Middle Ages, written documentation had worked its way into nearly every aspect of public and private life. Records that once traveled with a ruler's retinue solidified into fixed, permanent repositories, and there were many more of them: the Spanish crown's great central archive at Simancas, established between 1540 and 1561, alongside a proliferation of diplomatic, civic, guild, and ecclesiastical collections. The explosion of archives came, in early modern Europe, with a wider impulse to systematize. In the mid-seventeenth century the Czech reformer Jan Amos Comenius proposed to bring all knowledge into a single frame, developing a "pansophic" understanding of the world as one interdependent whole in the hope of fostering peace. The scientific dictionaries that proliferated during the early eighteenth century, as Richard Yeo has shown, took up the same ambition with a sharper sense of its limits, compressing and classifying the overwhelming flow of printed knowledge. Where the dictionary tried to contain knowledge, the encyclopedia tried to connect and circulate it: Chambers's Cyclopædia of 1728 gave up the dream of one stable map of the sciences for the pragmatism of the alphabet, and Diderot and d'Alembert carried the ambition to its famous height.

The frontispiece of Diderot and d'Alembert's Encyclopédie, engraved by Bonaventure-Louis Prévost, where Truth stands veiled and radiant beneath an Ionic temple while Reason and Philosophy tear away her veil, with the sciences and arts in ranks below

The frontispiece of Diderot and d'Alembert's Encyclopédie, engraved by Bonaventure-Louis Prévost, where Truth stands veiled and radiant beneath an Ionic temple while Reason and Philosophy tear away her veil, with the sciences and arts in ranks below. Source ↗.

The same Enlightenment drive, to systematize knowledge such that it would be accessible to any rational mind, was institutionalized at the turn of the nineteenth century in the archival state. In the 1790s revolutionary France turned the documents once held by the crown and church into the property of the nation. The archive in this way ceased to be the private memory of the state and became, in principle, the collective memory of the people. Though the shift from personal to institutional record-keeping was long and contested, as public officials continued to keep and sell their papers, the nineteenth century in general saw a turn toward more open record keeping. Britain, once the fear of revolution had passed, established its Public Record Office in 1838.

Max Weber would later place the archive at the center of his account of bureaucracy and rationalization and, indeed, the Victorian period saw the archive bound to the counting state. National statistical offices began systematizing social and economic data, while colonial administrations did the same on a global scale, gathering information about subject populations in the service of control. The British Empire was, for all its later air of grand design, something of an improvisation, run by an overstretched and undertrained service that compensated for its lack of men on the ground with information, surveying, mapping, counting, listing, and sorting the results into ever-shifting classifications.

In the twentieth century, the same impulse would be expressed in a variety of forms, as successive technological innovations seemed to herald the dawn of the universal archive. In 1895, Paul Otlet and Henri La Fontaine began the Universal Bibliographic Repertory, millions of index cards intended to link all global knowledge, which by 1910 had grown into the Mundaneum in Brussels. The microfilm boom of the 1930s, as Ian Milligan recounts, raised hopes of universal access and permanent preservation. UNESCO and other postwar institutions recast the imperial ambition to survey subject populations as global cooperation. In turn the advent of computing brought new possibilities. Ted Nelson coined the term "hypertext" in his effort to create Xanadu, a free, democratic library of linked information, while the cliometricians of the 1960s digitized census and economic records in bulk.

None of these projects ever became what their founders imagined, yet each achieved more than a cautious effort would have. It is no exaggeration to say that, in archival history, hubris is the method.

A plate from Paul Otlet's Atlas Monde, drawn in the 1930s for an Encyclopaedia Universalis Mundaneum that would have rendered the whole of knowledge in diagrams

A plate from Paul Otlet's Atlas Monde, drawn in the 1930s for an Encyclopaedia Universalis Mundaneum that would have rendered the whole of knowledge in diagrams. Source ↗.

Hubris as method

To set out to assemble the textual past of the whole world is, by any sober measure, absurd. A corpus so vast will never be comprehensively gathered by anyone. But such hubris is methodologically valuable. Where a project scaled to what is immediately reachable builds only the tools that task requires, an all-encompassing one must confront problems that a narrower effort would simply design around. The apparatus built to chase such an impossible goal can long outlast the dream that inspired it. Otlet's Mundaneum collapsed, but the Universal Decimal Classification he developed to organize it can still be found on library shelves around the world. Even if a universal archive is never assembled, what gets built along the way can dramatically expand the possibilities of historical research.

Still, some sobriety is in order. If the turn of the twenty-first century does herald the beginning of a more truly global archive, it is worth recalling just how many times that moment has already seemed to arrive. A crowdsourced archive on this scale would require an intermediary with an unusual combination of qualities to hold it together. It would need technical, linguistic, paleographic, and historical expertise, durable funding and stewardship, serious attention to rights, ethics, and data security, as well as the ability to keep archives, scholars, and communities coordinated over decades. It would also need the thing such efforts most often lack: an incentive structure that makes sustained participation worthwhile for the individual scholar and archivist, rather than relying on goodwill or the rhythm of grant cycles.

Our four-step vision

What makes the dream more plausible today than at any other point in history is that the most stubborn obstacles are finally giving way. Digitization and mass photography now make it possible to crowdsource images of archival material at low expense; machine transcription can turn those reproductions into usable text; and interoperable frameworks and semantic tagging promise to weave the results into a single, searchable whole.

Leo is moving toward this goal in four steps. First, we are coordinating the crowdsourced gathering of images of archival material, to break the most basic bottleneck in bringing the world's textual heritage online. Second, we are transforming those images into machine-readable text using our state-of-the-art transcription model, which gives researchers a reason to upload their materials in the first place. Third, we are building a platform for institutions, archivists, and historians to correct those transcriptions, collaborate on assembling collections, and reconcile metadata and description across incompatible standards. Fourth, we are planning the semantic, statistical, and visualization layers that will let scholars trace patterns of movement, exchange, and transformation across time and space.

Digitization and mass photography

The vast quantities of material that would constitute a universal archive must first of all be brought together. Digitization makes this possible, but since the advent of the Internet, only a tiny share of the world's archival heritage has been brought online. Traditional digitization is slow and expensive, requiring specialist departments, professional imaging rigs, high-resolution scans, and, for most of the cultural heritage sector, a reliance on public-private partnerships to foot the bill. A euphoric wave of digitization work in the 1990s and 2000s, much of it driven by improvements in optical character recognition for printed sources, brought vast quantities of books, pamphlets, and newspapers online, but it also produced the fragmented landscape we now inhabit. The result is a patchwork of vendors and genealogy companies whose holdings sit in walled-off silos, the majority of them behind paywalls.

The opening comes from the advance in portable photography over the past decade. The camera in an ordinary smartphone is now more than good enough to capture every character on a typical page, and many historians already hold large personal collections of such images, accumulated over years of archive visits. Simply by uploading images they already have, scholars can gradually build an enormous shared archive of primary sources.

Machine transcription

Though some altruistic scholars may digitize manuscripts purely out of generosity, most are only likely to do so with the prospect of getting something tangible in return. To incentivize users not just to upload their material for private use, but to publish it for the benefit of others, we have launched the Leo Transcription Grant, offering free transcription credits in return for releasing the images and corresponding transcriptions without copyright or restriction. 

The key, then, is to offer fast and accurate transcriptions for the widest possible range of source material. Our state-of-the-art model handles both printed and handwritten text without requiring users to train it themselves or to wrestle with preprocessing and segmentation. Its scope, for now, is limited to Latin scripts from roughly the past five hundred years, including English, French, Italian, Dutch, Spanish, and German, but not yet Greek, Arabic, Hebrew, Cyrillic, or Indic and East Asian scripts. The plan is to put the foundations in place before widening the scope. Our immediate priorities, on which everything else depends, are crowdsourcing collections and building the machinery to keep improving the transcription model. That is why we are focusing our energies on making Leo as user-friendly and powerful as possible, encouraging scholars to correct transcriptions within the web app and thereby generating a continuous flow of data that can be used to fine-tune the model.

Interoperability

Gathering and transcribing images is the foundation, but a heap of images and transcriptions is not yet an archive. The next step is the hard problem of imposing order on such an unwieldy mass of material.

Traditionally, the answer to the question of archival organization has been systematic arrangement according to some central schema. The cabinet improved upon the medieval chest by introducing doors and drawers that mirrored the divisions of knowledge. But where such approaches imposed a single source of order, under the physical constraint that a document could occupy only one position, digital reproductions allow for more flexibility. More than that, once an image has been transcribed, the text becomes a source of structure in its own right. A model can extract the names, dates, and places it contains, generate descriptive metadata, and align it with what other collections already hold.

A cabinet of hooks, each labelled with a subject heading, on which a scholar hung slips of paper excerpted from his reading. From Vincent Placcius, De arte excerpendi, 1689.

A cabinet of hooks, each labelled with a subject heading, on which a scholar hung slips of paper excerpted from his reading. From Vincent Placcius, De arte excerpendi, 1689. Source ↗.

What this makes possible is organization through interoperability rather than centralization, where separate collections, in different languages and formats, are made to speak to one another. Perhaps a better word, for humanists, would be translation. If Leo can act as such a mediating layer, it can become the connective tissue that Manning imagined for making global history possible.

In practice, rendering diverse collections interoperable is its own immense challenge, not least because the drive to standardize has itself bred a thicket of competing schemes. The Text Encoding Initiative (TEI), the standard for marking up archival text, was designed for interoperability, yet it is so verbose and so variously applied that one TEI file is often incompatible with the next, and any tool meant to work across corpora must first normalize them. One possible approach is to keep TEI at the center for archival fidelity while pairing it with lighter, machine-friendly formats for analysis and exchange, with linked-data ontologies such as CIDOC-CRM for the semantic web, and with IIIF for interoperable images, and then to map archival description standards onto a shared core rather than forcing every collection onto a single standard. Here too machine learning will play a critical role. The same kind of model that turns an image into text can translate between schemas, automating much of the mapping between incompatible standards that once had to be done by hand.

While such minutiae still need to be resolved, we see the path forward as a minimal shared core with automated integrations between standards, creating enough common ground to make collections legible to one another without forcing each to abandon its traditional indexing. Models for such an approach already exist, for instance in the Native Northeast Portal, formerly the Yale Indian Papers Project, which draws together primary sources relating to First Peoples from institutions across the northeastern United States, reorganizing them around Indigenous polities and concerns while preserving each document's reference to its original archival home. In doing so it repairs connections that earlier collecting practices had severed and helps reconcile the work of community historians with their professional counterparts.

The task presents most difficulties where resources are limited. There are already aggregators that normalize metadata across collections, including Archives Hub in Britain, the Digital Public Library of America and ArchiveGrid in the United States, and Europeana and Archives Portal Europe on the continent. But many regions have no comparable center of gravity. One task for a universal archive is to connect those aggregators where they exist and to stand in where none do.

Tagging

Finally, on top of all this sits the layer that makes the archive usable as evidence at scale: the standard apparatus of corpus linguistics, including frequency lists, keyness statistics, n-grams, collocations, and concordances, as well as semantic tools that extract entities and link them into knowledge graphs. This demands more engineering work than simply calling a large language model, because historians need tools that keep provenance in view. A navigable graph of evidence has to be traceable back to the transcriptions and original images, rather than serving up plausible answers whose workings are hidden.

Tagging is where this apparatus begins. To mark a span of text as a person, a place, a date, or a crime is not simply to label it but to make an interpretive claim, and the categories that matter are seldom given in advance. What one historian needs to identify—whether it is a kind of dwelling, a category of offense, or a form of kinship—another will have no use for, and a system that ships with a fixed vocabulary forecloses the possibilities that such an archive is intended to enable. A tagging feature should let scholars define their own categories rather than choose from a menu, and attach attributes beneath each tag. That way, a single person can be linked simultaneously to a stable set of facts, such as a date of birth or an occupation, while recurring across documents in different roles, for instance as a plaintiff in one and a defendant in another. Extracting persistent identifiers for people, places, and events means that the same person mentioned in a Flanders parish register might be traced again in a Bahia ship's log.

The components of such a system already exist, but no single tool yet brings them together. Transkribus lets users export tags together with their surrounding text. OCHRE, at the University of Chicago, stores material in a graph rather than a relational database, giving it the flexibility to model the tangled, many-to-many relationships that historical records contain. What none yet offers is meaningful assistance from AI tools. Every tag must still be placed by hand, which leaves tagging, like traditional approaches to imaging, transcribing, and cataloguing archival materials, impossible to accomplish at scale.

The long-term vision for Leo is thus not just to assemble, transcribe, and organize archival materials, but also to make sense of them. Here, even more than in transcription, the model cannot be left to work alone. Because a tag is an interpretation, the historian's judgment is a central part of the process. To work well, the model will require constant human input as feedback to sharpen its suggestions. Such an approach keeps every claim traceable to the transcription and image beneath it, holding the work to the discipline's standards of evidence, while relieving archival historians of the drudgery of sifting through countless sources by hand and freeing them for the interpretive work that only they can do.

How we're getting started

To propose assembling the textual past of the entire world may be immodest to the point of absurdity. The work will never be finished, but it is worth attempting anyway. It is in that spirit that we introduce Leo Collections. Our first public collection, ExLatinis, sets out to translate every Latin work printed in Europe between 1450 and 1750, a small but concrete start on a very long road. You can read more about that project here.

If you have images of archival material, apply for a Leo Transcription Grant and we will transcribe them in exchange for your releasing the images and transcriptions in the public domain. Or if you have another idea about how to get involved, get in touch at jon@tryleo.ai.

Share this article

© 2026 Leo Technologies Limited. All rights reserved