“The ISPS data archive aligns with the core mission of ISPS to uphold the very best practices in all aspects of social science research,” said Limor Peer, associate director for research and strategic initiatives at ISPS and senior research support specialist at DISSC. “And to back it up with resources.”
In 2018, ISPS launched YARD (Yale Application for Research Data), an open-source web application for reviewing and enhancing research outputs that feeds into the ISPS data archive.
ISPS staff work with social science researchers to review their files, documentation, data, and any code for studies prior to publication, as part of a process that has influenced and informed the development of similar practices in other social science data archives that recently joined together under a consortium called Curating for Reproducibility (CURE).
“We make sure that people can use the digital material as intended,” Peer said. “If there’s a data file, we make sure they can open it and understand what’s in it. If there’s a script, we ensure they can run it, and it works. We want the data and code to be as usable as possible to the scientific community for as long as possible, so the community can do science.”
But outside of ISPS and its affiliates, faculty members have often found their own storage solutions to meet current standards for data management, reproducibility, and greater ease of sharing increasingly demanded by research funders and academic journals.
“When researchers submit grant proposals, they need to be able to describe how they plan to manage their data and ensure that it be safe and accessible for the long term,” said Rebecca Dikow, director of computational methods and data at Yale Library. “This has become a growing burden that Yale Dataverse can alleviate.”
In 2019, ISPS began to meet with the library to build what would become Yale Dataverse, based on a platform created at Harvard University and eventually made available for Yale’s use.
In 2022, DISSC arrived on campus and joined the effort, having emerged from a series of recommendations (PDF) by a committee of social scientists from across the university. DISSC now serves as a campus hub to manage and facilitate the collection, protection, and utilization of new, frequently large datasets that are currently revolutionizing fields like political science, economics, psychology, and sociology.
“This has been more than a seven-year journey to bring Yale Dataverse to fruition,” Peer said. “We are incredibly grateful to our peers at Harvard and our partners across the university.”
The Yale Dataverse team explored the metadata needs for Yale researchers to efficiently search for files, piloted the process of uploading data with select faculty members, demonstrated the capacity to harvest data from other repositories, documented the process to create non-Yale collaborator accounts, and worked with faculty members to test the system.
“We have intentionally broken this many times to ensure its robustness,” said Barbara Esty, head of data services for Yale Library, noting the disparate, uncoordinated methods that researchers currently use to manage their data, such as websites and external hard drives. “We are trying to create a system that works for us, using what we have learned from our experience elsewhere.”
Joshua Kalla, ISPS faculty fellow and associate professor of political science, said that as a researcher, he values how this new repository enhances research transparency and strengthens the reproducibility of scholarly work.
“The ability to easily archive, discover, and access research data through Yale Dataverse will help ensure our findings can be properly replicated and built upon by other scholars, expanding our research’s impact and reach,” Kalla said.
ISPS Director Alan Gerber, Sterling Professor of Political Science, sees the repository as an extension of rigorous scientific work.
“There is a growing expectation that researchers make the data underlying their work available in a reliable and easy-to-use form,” Gerber said. “It is great to see how the Yale Dataverse is dramatically expanding what ISPS has done to build tools and infrastructure in support of this collective scientific project.”
The library aims to recruit a dedicated research data management librarian for the Dataverse, improve the interface, and integrate other systems like YARD.
“My goal would be for this service to get to a point where once the researcher has completed their research and deposited their data, they don’t need to worry about it,” said Dale Hendrickson, senior director for library information technology. “What better place than the library to play a role in that, since that’s what we do.”