José-Antonio Espín-Sánchez found the collection of 15 books in a library of the Instituto Hispano Cubano, in the Spanish city of Seville.
An associate professor of economics, Espín-Sánchez has been collecting data on all immigrants from Spain to the Americas since the time of Christopher Columbus and linking their family trees. One 16th century collection he sought contained contracts with ship captains and passenger listings. Yale only had the 10th volume, and Harvard none. The collection had never before been scanned. No intelligence — artificial or otherwise — had likely attempted to examine the text and extract names, dates, and relationships for large-scale analysis.
With the help of ChatGPT, a generative artificial intelligence large language model (LLM), Espín-Sánchez reduced the time needed to transcribe and process each record from hours to minutes, potentially reducing the overall project time from decades to years.
“The idea is that when we have the whole dataset, we can create family trees for the last 500 years,” he said.
But with this new and rapidly evolving technology, researchers must take great care to ensure accuracy and overcome obstacles. At a recent collaborative session sponsored by Yale’s new Data-Intensive Social Science Center (DISSC), the Institution for Social and Policy Studies (ISPS), and the Institute for the Foundations of Data Science (FDS), Espín-Sánchez joined four other social scientists and a room of researchers to share what they are learning about how AI can help them learn.
Kevin DeLuca, an ISPS faculty fellow and assistant professor of political science, discussed how he uses ChatGPT to collect and analyze historical newspaper endorsements of political candidates. The AI can help fix errors from more established optical character recognition (OCR) applications and place the data into a structured format, significantly reducing the time and effort required to manually transcribe text from newspaper archives and organize data to more easily answer research questions.
“It spits out a table usually within a couple of seconds, which is pretty great,” DeLuca said.
As with Espín-Sánchez’s examination of scanned documents, the AI can struggle with accuracy, often due to OCR errors or formatting issues. For example, scanning old bound books face down often introduces a curve to the lines of text because of how the pages are attached to the spine. And the software can lose its place when following a curve from one line to another across the page, jumping around, breaking up names, and changing the order of the words.
But the researchers have found effective ways to mitigate these problems. And they have begun evaluating the different available AI tools to fine-tune the best approaches for their tasks.
Emma Zang, an ISPS faculty fellow and assistant professor of sociology, biostatistics, and global affairs, used AI to decode 1.5 million legal judgments in China to study the effect of a 2011 change in law on gender inequality in divorce cases. The AI extracted and processed vast amounts of data from 2009 to 2016 — written in Chinese and involving gender, child custody decisions, and housing property disputes — that would have been impractical to analyze manually.