Skip to main navigation Skip to search Skip to main content

Predicting the replicability of social and behavioural science claims in COVID-19 preprints

  • Alexandru Marcoci
  • , David P. Wilkinson
  • , Ans Vercammen
  • , Bonnie C. Wintle
  • , Anna Lou Abatayo
  • , Ernest Baskin
  • , Henk Berkman
  • , Erin M. Buchanan
  • , Sara Capitán
  • , Tabaré Capitán
  • , Ginny Chan
  • , Kent Jason G. Cheng
  • , Tom Coupé
  • , Sarah Dryhurst
  • , Jianhua Duan
  • , John E. Edlund
  • , Timothy M. Errington
  • , Anna Fedor
  • , Fiona Fidler
  • , James G. Field
  • Nicholas Fox, Hannah Fraser, Alexandra L. J. Freeman, Anca Hanea, Felix Holzmeister, Sanghyun Hong, Raquel Huggins, Nick Huntington-Klein, Magnus Johannesson, Angela M. Jones, Hansika Kapoor, John Kerr, Melissa Kline Struhl, Marta Kołczyńska, Yang Liu, Zachary Loomas, Brianna Luis, Esteban Méndez, Olivia Miske, Fallon Mody, Carolin Nast, Brian A. Nosek, E. Simon Parsons, Thomas Pfeiffer, W. Robert Reed, Jon Roozenbeek, Alexa R. Schlyfestone, Claudia R. Schneider, Andrew Soh, Zhongchen Song, Anirudh Tagat, Melba Tutor, Andrew H. Tyner, Karolina Urbanska, Sander van der Linden

Research output: Contribution to JournalArticleAcademicpeer-review

Abstract

Replications are important for assessing the reliability of published findings. However, they are costly, and it is infeasible to replicate everything. Accurate, fast, lower-cost alternatives such as eliciting predictions could accelerate assessment for rapid policy implementation in a crisis and help guide a more efficient allocation of scarce replication resources. We elicited judgements from participants on 100 claims from preprints about an emerging area of research (COVID-19 pandemic) using an interactive structured elicitation protocol, and we conducted 29 new high-powered replications. After interacting with their peers, participant groups with lower task expertise (‘beginners’) updated their estimates and confidence in their judgements significantly more than groups with greater task expertise (‘experienced’). For experienced individuals, the average accuracy was 0.57 (95% CI: [0.53, 0.61]) after interaction, and they correctly classified 61% of claims; beginners’ average accuracy was 0.58 (95% CI: [0.54, 0.62]), correctly classifying 69% of claims. The difference in accuracy between groups was not statistically significant and their judgements on the full set of claims were correlated (r(98) = 0.48, P < 0.001). These results suggest that both beginners and more-experienced participants using a structured process have some ability to make better-than-chance predictions about the reliability of ‘fast science’ under conditions of high uncertainty. However, given the importance of such assessments for making evidence-based critical decisions in a crisis, more research is required to understand who the right experts in forecasting replicability are and how their judgements ought to be elicited.
Original languageEnglish
Pages (from-to)287-304
JournalNature Human Behaviour
Volume9
Issue number2
DOIs
Publication statusPublished - 1 Feb 2025
Externally publishedYes

Funding

This research was developed with funding from the Defense Advanced Research Projects Agency (DARPA) under cooperative agreement nos. HR001118S0047 (A.M., H.F., A.H., A.V., F.M., B.C.W., F.F., D.P.W.), HR00112020015 and N660011924015 (A.L.A., M.K.S., N.F., O.M., Z.L., E.S.P., A.H.T., B.L., T.M.E., B.A.N.) and N66001-19-C-4014 (T.P., F.H., M.J., Y.L.). Following the DARPA model, the funder played a role in conceptualization and design, but was not involved in data collection, analysis, decision to publish or preparation of the manuscript. The views, opinions and/or findings expressed are those of the authors and should not be interpreted as representing the official views or policies of the Department of Defense or the US Government. T.M.E., N.F., Z.L., B.L., O.M., B.N. and A.H.T. thank M. v. Assen, R. v. Aert, M. Bakker and M. Sitnikov for help with power calculations and calculation of effect sizes; V. Ashok for contributions to the replication outcome data; and the hundreds of researchers who conducted the replications, served as editors and reviewers for the preregistration review process, and helped conduct power analyses and additional statistical consulting. A.M. thanks D. Siegel for research assistance, and the audience of the SKAPE seminar at the University of Edinburgh for helpful feedback on an earlier version of this paper.

FundersFunder number
U.S. Department of Defense
US government
Defense Advanced Research Projects AgencyN660011924015, HR001118S0047, N66001-19-C-4014, HR00112020015

    Fingerprint

    Dive into the research topics of 'Predicting the replicability of social and behavioural science claims in COVID-19 preprints'. Together they form a unique fingerprint.

    Cite this