Good science requires more than a good experiment. It invites others to assess the work and to reproduce and replicate the results. To do so, researchers need to provide access to their methods and data.
Facing growing pressure from funders and evolving journal standards, the Institution for Social and Policy Studies (ISPS), the Data-Intensive Social Science Center (DISSC), and Yale Library are holding a series of training sessions for faculty, students, and staff about open and reproducible research (ORR).
“These sessions are not just about compliance,” said Limor Peer, associate director for research and strategic initiatives at ISPS and ORR program lead at DISSC. “They are about teaching scholars how to share their research with the same rigor they apply to conducting their research. Attendees come away with practical tools, guidelines, and a sense of community around the shared challenge of making science more transparent and trustworthy.”
Last week, Peer launched the series with a session on who reproducible research is for and why it matters. She began by clarifying the differences between reproducibility (the ability to produce the same results with the same data and the same code), replicability (the ability to reach similar conclusions using new data and independent methods), and robustness (the degree to which results hold under different assumptions, models, or analytical choices).
Reproducibility aims to strengthen scientific results by enabling verification and generalization of results by peers.
“Really what we want to be able to do is codify, expand, and instill practices that follow scientific principles, which include investigating, testing, and self-correcting when warranted,” Peer said. “These practices will also improve your productivity, allow you to verify your own results, and enable other people to extend the work.”
Anthony Lollo, director of data science and analytics at Yale School of Public Health, and Maurice Dalton, DISSC’s data engineering and solutions lead, also contributed to the first session.
Lollo presented a detailed walkthrough of his team’s journey preparing a replication package for their paper on Medicaid privatization in Louisiana. The project spanned seven years, involved data limited by federal privacy laws protecting medical information, and required coordination across multiple researchers and programming languages.