Replication does not happen at scale, in part, because of a market failure. Strong incentives drive researchers to publish new findings, with funding and prestige both flowing from novel claims. On the other hand, verification carries weak incentives; little glory comes from confirming someone else’s work. Previous attempts at large-scale replication have failed because they required armies of specialists to verify each study by hand. Manual verification cannot scale to the millions of papers published annually, and the problem is about to grow far more acute.
As AI is introduced into the scientific process, it risks compounding these problems. False findings will multiply as it becomes easier to generate plausible-sounding scientific results than to verify them. AI research offers a preview. Leading conferences have seen submission surges of 60% in a single year, overwhelming the field’s capacity to evaluate new results. Researchers are now burdened with reviewing nonsensical AI-generated submissions while rebutting low-quality AI-generated reviews of their own work.175 Other fields will follow the same trajectory.
The Genesis Mission is building a science generator with instruments capable of producing scientific discovery at an unprecedented scale. To sustain progress, we must also build its necessary counterpart: a verifier equal in rigor and scale. This is the central challenge that must be undertaken to address the reproducibility crisis and capture the full benefits of AI for science.
AI itself could help close the generation-verification gap, but only if we invest in the necessary infrastructure. AI has already begun to automate significant parts of the scientific workflow. Meanwhile, the Gold Standard Science requirements, including reproducibility, data sharing, and methodological documentation, create precisely the conditions under which automated verification becomes possible. The combination of both could lead to low-cost, continuous AI-enabled verification. Researchers have already outlined one vision of such a system, in which specialized agents parse submitted papers, reconstruct computational environments, execute analyses in sandboxed settings, and compare outputs against claimed results.176 The same infrastructure that audits human-authored papers today could tomorrow judge which machine-generated hypotheses merit experimental resources.
Rising to this moment of need, NIH has launched a new, agency-wide initiative to elevate replication and reproducibility studies, identifying critical research and infrastructure needs to advance rigorous findings that are verifiable and transparently shared. Looking forward, the Federal Government must continue to lay the connective tissue between verification infrastructure and our scientific enterprise. This means establishing open APIs and interoperability standards that allow verification capabilities to plug into journal submission systems, grant reporting platforms, and private-sector AI research tools; standards for replication packages that ensure computational research arrives in machine-auditable form; and prizes for successfully replicating or disproving influential papers. The result should be a verification system that is not occasional but continuous, low-cost, and commensurate with the scale of discovery we are now capable of producing