“We Have Returned to Dreaming Big”: Kirt von Daacke on How AI Transcription Is Helping to Recover the Early History of the University of Virginia
Kirt von Daacke has spent nearly fifteen years orchestrating Jefferson's University — The Early Life (JUEL), a digital archive of the University of Virginia's first half-century assembled from ledgers, letters, diaries, and court records, most of them handwritten. For over a decade, students transcribed those documents by hand, one page at a time. Automated transcription removed that bottleneck, and the project has since pushed back to the 1770s, outward into the surrounding county, and into civil court suits that can run past a thousand pages. We spoke with him about what changes when the reading stops being the constraint, and about why a historian must still be in the room.
Jon Cooper
Founder and CEO · August 18, 2026

Contents
Kirt von Daacke is a Research Professor at the University of Virginia, where he teaches in History and American Studies. His research and writing focus on slavery and race in the U.S. South. He co-founded Jefferson's University — The Early Life (JUEL) and is Executive Director of the Gibbons Project, the Provost's Office initiative documenting those enslaved at UVA. He co-chaired the President's Commission on Slavery and the University as well as the President's Commission on the University in the Age of Segregation, and serves as Managing Director of Universities Studying Slavery, a consortium of more than a hundred institutions across six countries. He is the author of Freedom Has a Face: Race, Identity, and Community in Jefferson's Virginia (2012) and a co-editor, with Andrea Douglas, of After Emancipation: Racism and Resistance at the University of Virginia (2024). JUEL is a recipient of a Leo Transcription Grant.

“Die Virginia-Universität in den vereinigten Staaten von Nordamerika,” Johann Poppel, 1837: the Academical Village roughly midway through the period JUEL covers. Public domain, via Wikimedia Commons
Recovering the early history of the University of Virginia
Could you introduce yourself and describe the JUEL project? How did the project get started, who it's for, and what it's trying to recover?
I’m a history professor at the University of Virginia, my research and writing have focused on slavery and race in the U.S. South. Fourteen to fifteen years ago, while teaching a class on slave testimony where we were doing close textual readings of written materials produced by enslaved people in the United States, a colleague in art history who studied the same period suggested we take my class on a tour of UVA. I loved the idea, we needed to step away from the textual analysis and think about space and place, and how enslaved people inhabited and survived in those spaces. It was a great day. At the end, one student asked us: “How do you know all this about UVA?” We pointed to the library and indicated that it was all just sitting in the university archives. The student then asked, is anyone doing anything with it? We looked at each other and realized that we should be doing something with it.
Neither of us had any digital humanities training, but as we thought about what we might do, we came to two sets of desires: First, that the documents would be accessible and visible; they would have metadata tagging; and we’d be able to track people across document streams and places and begin to map human activity and networks (this was my particular interest, I have used network theory in my earlier work). Second, we wanted to create a complex 3D model of UVA that would resurface the ways in which it was a complicated landscape of slavery. Where did the enslaved live? What work did they do? How did Jefferson think about their presence in his design for a university? How did slavery and the enslaved come to shape the university? We wanted a digital humanities project that would do both and integrate them into a textual and visual whole that allowed for actual research. As we thought about the documents, though, we realized that they could tell many other stories about the development of the first elective curriculum in U.S. higher education, about the revolution in college design UVA kicked off, about how the school grew and changed. It was a complex and powerful window into American history. Those conversations, once we connected with another professor who designed digital humanities projects, soon became Jefferson’s University-The Early Life project. We imagined it as a learning and research site that could be used by scholars, college students, alumni, and even secondary school students. Really, we wanted it to be a public good making a rich archive of documents accessible and understandable that would tell the full and messy story of the university’s first fifty years. It would include official university records as well as personal papers of those who lived and worked there, and it would be built so it could expand in two ways. One was outward into the surrounding community, as the university was never separate from its region. We would be able to include details about life in a college town and a rural plantation county or counties, and eventually much farther afield than that. The other was that although it was initially confined to the university’s first fifty years, we wanted it to be built so that it could incorporate more recent material and also tell stories from the next one hundred and fifty years of school history.
From student transcription to automated text recognition
The material you're dealing with is unruly. It runs from the mid-eighteenth to the later nineteenth century and includes ledgers and university records, letters, journals, diaries, court records, business records, alongside printed speeches and essays - mostly handwritten and captured in a variety of formats, from phone photographs snapped in a reading room to microfilm to library-quality digital scans. Could you describe what the project's workflow looked like before AI transcription entered it? How has being able to process these records at scale changed the scope of the project?
Fascinatingly, although we anticipated the project eventually moving beyond the 1870s and toward, if not into, the 20th century, in this current phase of development, we’ve actually gone backwards in time, incorporating local and regional material from before the university was founded. We now include information from the 1770s onward. When we started the project in 2011-2012, AI was the stuff of science fiction, so we did three things: First, for many items, we simply had researchers go to special collections and do the transcription themselves or there were existing typescript or published transcriptions that we could immediately use. Wherever possible, we used those and the JUEL website would simply showcase a transcription that had been tagged. Second, if we had handwritten material we needed scanned that was far too cumbersome to sit in special collections and transcribe, we asked the library to create digital scans. That worked for a few collections, but library scanning was then quite expensive. Third, when working with existing scans, we either farmed those out to transcription services (which were done by humans—it was expensive and frankly the product was not that great, it required quite a bit of human cleanup on our end) OR we assigned them to students whom we had trained in documentary transcription. They literally transcribed the documents by hand, then did the metadata tagging (in Oxygen, they were learning XML on the fly). We wanted document scans and transcriptions to be visible side by side, but could not then make it work. It was slow—we had students doing all the transcription for over a decade—they did a lot of great work, but it was entirely on items written by people living or working at UVA—we had not yet left the campus, and that took a decade to complete enough material to feel like we were telling new and complicated stories about the school’s development. By 2017, phone technology had improved to the point where we could easily do our own scanning, using an app to take the pictures, save each document as a PDF file, and import to the cloud. We suddenly were generating a huge backlog of scanned documents in need of transcription, but we did not have the ability to move quickly. Thus, the expansion, both in geographic area and the temporal, awaited a future when we had more help and more money.
In 2024, we received a grant to hire students anew and return to transcription, but things had changed. Students no longer wrote or read cursive and found the work way too difficult, we struggled to keep them on payroll. One student asked the question: why aren’t we using technology to do some of this? I did not have a good answer, so charged him with figuring it out. He tried a number of software packages, all had what were unacceptably low rates of accuracy on transcription. That began to change in November 2024, when he discovered that ChatGPT 4 Omega had come out earlier in the year and demonstrated far greature accuracy. Suddenly, AI had our interest, but we had no idea how to do it. At his spring research presentation, another history professor turned to me and said—have you heard about Leo? I just met with them. That connection proved to be an immediate game changer—we could use Leo to do primary transcription at scale and were able to generate thousands of transcriptions and ingest them into JUEL, where students would do the work of secondary transcription (cleaning up and correcting the AI transcription), which they could handle when they had eighty to ninety-five percent of the document accurately transcribed already. They could also do tagging at a faster clip. We were then able to start collecting a much more diverse and wide-ranging set of documents that had more recently been scanned (HathiTrust and libraries everywhere had leaned into digitizing documents since 2016). We have returned to dreaming big about what we can ingest—we have already expanded the geographic scope and have material ready for expanding the temporal scope in the future.
What ATR-1 changed for complex historical documents
The release of ATR-1 changed what Leo could do: complex mixed-format tables, pages with text running in more than one direction, and the civil court suits with varied formats that the project has since moved into. Is there a common thread among the documents that used to be set aside and are now machine-readable, and what did having them open up? And which of your materials does automated transcription still struggle with?
We have thousands of documents that are complex and mixed format and AI had traditionally struggled with tabular data, cross-writing, additional notes running in different directions on a page, and documents with highly varied formatting. This is especially the case in the civil court suits (collections of papers from each suit), which can run from three to more than 1,500 pages of mixed material.
Before ATR-1, we spent a lot of time reformatting tabular data so it at least appeared in semi-tabular form in the transcription, and the mixed format documents often required page-by-page re-transcription, often manually. ATR-1 changed all of that and happily forced our software development team to rewrite JUEL code so it could accommodate the new tabular data. Automated transcription now largely handles these complex and messy documents, and now typically struggles with the kinds of things that experienced humans struggle with: smudgy ink, non-standard abbreviations, pages in which writing seemingly runs in every direction.

A page of Albemarle County property tax records, 1797. Ruled columns, abbreviated headings, and a running tally of people counted as property. Public domain, via Wikimedia Commons
Open access through the Leo Transcription Grant
By your own account you've transcribed only a small fraction of the material under the project's scope, and you keep finding new material. JUEL received a Leo Transcription Grant toward that work, on the condition that the transcriptions and images be published openly and free of restriction. How has the grant helped, and what did it let you attempt earlier than you could have on your own? And how does that condition sit with you: is publishing everything openly a price of the arrangement, or something the project would have done regardless?
We could not have proceeded at this speed or anything close to it without the grant. This is a difficult time for digital humanities projects, especially those that often uncover difficult histories. Federal grants have largely ceased to exist, large scale private foundations have stepped back from funding these university projects, state level humanities councils have been hard-hit by federal cuts, and even internal university resources have proved far more scarce since 2022 (due both to fiscal contraction during pandemic and to post-pandemic federal attacks on higher education). We managed to secure private grant-funding to pay students to do research and transcription and to overhaul JUEL (we moved it to a new platform, redesigned the interface, etc.), but the project has almost no other funding, so paying for automated transcription has meant we could not tackle large document collections (again, civil court cases, where we have literally thousands of them and counting, and some individual suits run well over 1,000 pages). We are incredibly grateful for the support. Happily, the requirement that the material (both digitized documents and their transcriptions) be published openly is entirely in line with our project—it from the start has been about making the historical record visible, accessible, and intelligible to anyone with an internet connection.
That exchange - free transcription in return for public release - is the engine behind Leo Collections, our attempt to build a crowdsourced, machine-readable archive of the world's textual past, largely out of the images scholars already hold. You're doing a version of this yourself, bringing nineteenth-century material from different institutions onto a shared platform. What does a project stand to gain by joining its material to everyone else's, and what makes that hard in practice?
For our project, which has always understood itself as telling one university’s story as part of a connected and expanding universe (university and town, university and region, university and state, university and United States, university as part of a network of schools, and even sometimes the university as part of a global story), we have always imagined that it would eventually connect to other schools and other projects. If we are going to begin to tell such an expansive story, we have to join our material to everyone else’s. It is hard in practice because projects like mine develop in siloed fashion. There has not been an industry standard for data collection, for data preservation, for transcription, for the software undergirding the project, and often even for how one presents the material. Some digital projects that on the surface resemble JUEL have a public face that is largely explanatory and interpretive; some are simply forms of digital archive without extensive interpretation or contextualization. Sometimes projects like mine result in just spreadsheets recording the details found in documents, not the documents themselves. And, finally, the scholars building and driving these projects often don’t have that larger world in mind, nor do they have the bandwidth to even consider it.
From transcription to historical knowledge
A transcription on its own isn't yet history. JUEL is an archive of scans and transcriptions, as well as a relational database. Somewhere between the transcription and the final result, someone has to pull people, places, and events out of the text and connect them to one another. Could you walk us through what that actually looks like in practice? How do you get from thousands of pages of transcribed ledgers and court records to a structure a user can actually use?
We start with metadata tagging—when students have finished transcribing, our website software does initial metadata tagging, identifying people, places, and dates. Students have to narrow that down, tagging a person as the exact person already in the database or adding a new person to it; same goes for places. We also have them tag events and roles—we have a complex event hierarchy that captures most of the administrative features of life at a university as well as all the details about slavery. And every person has associated roles (student, faculty, merchant, enslaved person, etc.) The user interface for the website allows anyone visiting to search by a specific name, specific event, specific place, or by general categories. Each student working for us is required to once a year present at a research symposium and publish interpretive materials or essays on the website. For that, they develop a research project related to the material they have been transcribing and start to make the connections visible. The essays already published on the website demonstrate how that works and will continue to work. But we also know, especially since we have over thirty thousand unique individuals already identified in the system, that it will take a broader public searching, researching, and connecting. Thus, we ultimately see JUEL as a project that will continuously evolve as users and our own students make discoveries and connections.
Why a historian must remain in the room
What can a tool like Leo not do? Where in your workflow does a historian have to be in the room? A search returns only what you thought to ask for, which risks turning an archive into a device for confirming what you already believed, and the fact that a machine found nothing is never evidence that nothing is there. How do you guard against that, and how do you preserve the capacity to be surprised by a document that doesn't fit the categories you brought to it?
Yes, we have a historian in the virtual room as soon as students start secondary transcription. We use a messaging platform where we have a large group chat where all students share images and ask questions. This past year has been fantastic, with so many students working through so many documents, the conversations there are richer. Each student isn’t bogged down in a single letter for weeks at a time, so they much more quickly find connections, and we the scholars running the project are active listeners and teachers in those chats, helping with difficult sections that automated transcription did not quite capture or make mistakes on, helping with local names, familial connections, legal terms, you name it. We link students to reliable third party resources to help them answer questions on their own and have them use JUEL to try to figure things out as they go, but the complex search function at JUEL is public-facing and allows any user anywhere to discover new material. We even have a visible place where one can reach out to the JUEL team regarding mistakes, or even better, discoveries they’ve made. Yes, our students do interpretation by writing essays, but those cover a wide range of themes and topics and will never create a sort of party line, nor capture every story lurking in the expanding archive of materials. We do not prepare them for what they will find, we simply wait for them to ask questions and if they’ve found a productive and researchable topic, we cut them loose, as we are committed to telling as many stories as possible. I’ve always imagined this project as a digital example of Geertzian thick description. But this isn’t just about the students who work on the project. Classes and scholars across the university already use JUEL as a research tool, and those classes approach the history from vastly different angles.

Isabella Gibbons, enslaved by University of Virginia professors and, after emancipation, a teacher of freed African Americans in Charlottesville. Her words are inscribed on the university's Memorial to Enslaved Laborers. Public domain, via Wikimedia Commons
Public history, descendants, and dreaming bigger
JUEL is public history: not a dataset for specialists but a place where anyone can encounter what life at the early university was actually like, and where a descendant might look up a name. It's due to become one of three interrelated projects drawing on the same foundation, including one importing nineteenth-century data from other institutions and one focusing on the enslaved, with a three-dimensional reconstruction of the 1850s university alongside it. What changes for the public when material of this kind becomes easily accessible, and how does automated transcription make that possible?
We have always imagined, since 2011-2012, when the original JUEL project was born, that it would serve both a general public and to a degree, scholars. The project was imagined from the start as ultimately being so data rich that it would be of use to specialists, and that has been a reality. Other scholarly projects examining UVA or Virginia history have regularly used it and collaborated with us. Even better, it has already become a tool for a broader public seeking to learn about the university or learn about life, labor, and learning there. As we have updated and expanded the JUEL project, however, I think we are on the cusp of something truly significant. With automated transcription, our ability to vastly increase the types and quantity of material made easily accessible, we have been able to begin to connect a rich but local and particular project to a much wider world of inquiry. One of the project’s goals among many has been to lift up the names, lives, and stories of the university’s unsung founders, especially the enslaved. As our project expands into new archival ground, collecting court records from the surrounding area, we see it as not only identifying far more enslaved people, but identifying families and communities that stretch across plantations and counties. The goal is as well to move forward in time, past 1870, to include more and more material that will allow descendants of those enslaved in the region to trace the documentary record back to UVA or the area. The project does far more than sites like Ancestry and FamilySearch, as the material is transcribed, has metadata tagging that goes far beyond names and document types. It also provides opportunities for deeper interpretation.
The project has also expanded by thinking about UVA as but one school in a network of educational institutions in nineteenth century America that were deeply interconnected, sharing faculty, students, and ideas. We have also launched a third interrelated digital project (all three access the same underlying digital humanities architecture) that is mapping the human and intellectual networks comprising those schools. For the first time, all of the students, alumni, faculty, trustees, and administrators from dozens of schools will be identified and tracked in a single place. As well, the project tracks movement of individuals between schools and across state lines and as much as possible, collects their intellectual output (books, pamphlets, speeches, and the like).
Together, the three projects, through the shared data model, will help anyone trace ancestry as they learn about their own family histories. It allows users to examine documents, do their own research, and even share findings for publication on the website. Accessibility is the key here—we are collecting documents that are typically not a part of those ancestry sites’ collections—and transcribing them while making them deeply searchable. Automated transcription has allowed us to dream big, as we can now move through far more material far more quickly than before. That change cannot be overestimated—our production queue (from document scan through a three stage documentary transcription process and then metadata tagging) continues to grow, but Leo’s automated process and its high first pass accuracy rate even with complex documents with mixed formatting have made it possible for us to move much faster as we make documents accessible, intelligible, and visibly interconnected. This is a public history project built through a digital humanities framework that offer an interested public the tools to research enslavement, ancestry, labor, life, and learning in the world of nineteenth century institutions of secondary education.
Explore the project at Jefferson's University — The Early Life. Applications for Leo Transcription Grants are open to projects willing to publish their transcriptions and images openly.