Reproducibility is the capacity to obtain consistent results when research procedures are repeated under specified conditions. It supports the scientific method by making the relationship between evidence, methods, and conclusions independently examinable. Its meaning varies across disciplines. In computational research, reproducibility commonly means obtaining consistent results from the original data, code, and analytical procedures; repeating an experiment with newly collected data is usually distinguished as replication. Neither establishes, by itself, that a conclusion is correct or applicable beyond the original setting. (nationalacademies.org)
Terminology and scope
The 2019 National Academies report Reproducibility and Replicability in Science explicitly distinguishes computational reproducibility from replicability. Reproducibility concerns the same input data and computational methods, whereas replicability concerns independently collected data addressing the same scientific question. Generalizability concerns whether findings apply to different populations or contexts. These distinctions separate checking an analysis from testing whether its empirical findings recur. (nationalacademies.org)
A further distinction is repeatability, which generally concerns repetition under closely maintained conditions. In metrology, reproducibility conditions instead include changes in locations, operators, or measuring systems. The International Vocabulary of Metrology specifies that changed and unchanged conditions should be identified. Consequently, “reproducible” has no fully informative meaning without a statement of the procedures, conditions, and comparison criteria involved. (jcgm.bipm.org)
Terminology has also changed within computer science. In 2020, the Association for Computing Machinery revised its artifact-review terminology to align more closely with broader scientific usage. Its distinction between reproduced and replicated results depends on whether an independent investigation uses artifacts supplied by the original authors. (prod-www.acm.bloomreach.cloud)
Computational reproducibility
A computational result depends on more than its published formula or named algorithm. Relevant components include input files, preprocessing, parameter settings, source code, software dependencies, and execution conditions. A methods section may therefore be insufficient to reproduce a complex analysis. Missing files, undocumented manual operations, ambiguous instructions, and obsolete software can prevent reconstruction even when the reported method is scientifically reasonable. (nationalacademies.org)
Reproducibility does not always require identical binary output. Exact agreement may be appropriate for deterministic calculations, but some investigations accept differences within a stated numerical tolerance. Changes in computational environments and arithmetic behavior can affect output; the scientifically important question is whether those differences alter the reported results or conclusions. Comparison criteria should therefore distinguish exact computational identity from an acceptable range of variation. (nationalacademies.org)
Re-executing an analysis also does not establish the adequacy of its assumptions. The same procedure can consistently reproduce a programming error or an inappropriate application of statistics. Computational reproducibility makes the analytical process inspectable; evaluation of experimental design, interpretation, and supporting evidence remains a separate task. (nationalacademies.org)
Documentation and research infrastructure
Reproducible research preserves a traceable path from inputs to outputs. Data provenance records where data originated and how they were transformed. Version control identifies particular revisions of code and data, while workflow-management systems describe dependencies and execution order. Together, these mechanisms help associate published tables and figures with the procedures that generated them. (nationalacademies.org)
Useful research packages include executable code, clearly identified data, parameter configurations, and software documentation. Archival repositories and persistent identifiers help preserve and locate these materials. Sharing open-source software facilitates inspection, but accessibility alone does not establish that a package is complete, executable, or sufficient to reproduce its associated publication. (nationalacademies.org)
The FAIR data principles, formally published in 2016, describe research objects as findable, accessible, interoperable, and reusable. They emphasize machine-actionable metadata and apply to analytical tools and workflows as well as datasets. FAIR accessibility can include authentication and authorization: it is not synonymous with unrestricted public access. This distinction matters where data privacy or other restrictions limit distribution. (nature.com)
Machine learning
In machine learning, documentation must cover both model construction and evaluation. Relevant information includes training data, preprocessing, model architecture, hyperparameters, optimization procedures, and performance measures. The NeurIPS 2019 reproducibility program combined a reporting checklist, code-submission policies, and an independent reproducibility challenge to improve how research methods and results were communicated and assessed. (arxiv.org)
Software-level determinism is one component of this problem. Frameworks can provide deterministic implementations of operations, but enabling them does not, by itself, guarantee that an entire application is reproducible. The surrounding data-processing and execution pipeline also matters. Distinguishing a reproducible implementation from convincing evidence for a model’s claimed performance prevents these different questions from being conflated. (docs.pytorch.org)
Assessment and scientific interpretation
Independent checks may test whether research artifacts are available, whether they function as documented, or whether the main results can actually be obtained. ACM’s badging system separates artifact availability, artifact evaluation, and results validation. These categories recognize that providing materials and successfully reproducing a result are different achievements, rather than interchangeable forms of peer review. (prod-www.acm.bloomreach.cloud)
Replication assessment considers uncertainty rather than requiring identical observations. Effect estimates, confidence intervals, and the interpretation of statistical hypothesis tests can inform comparisons, but no universal criterion applies to every research field. A single failed replication does not conclusively disprove an original hypothesis, just as one successful replication does not guarantee correctness. Differences can arise from chance, uncontrolled conditions, methodological weaknesses, or previously unrecognized features of the system under study. (nationalacademies.org)