Skip to main content

Blog

Replicability vs Reproducibility: Check the Evidence

Understand replicability vs reproducibility by checking whether a report reuses original data or collects new data. Keep methods, results and gaps visible.

Semantic Map: Visualize the topic from new angles.
Knowledge Map: Deconstruct the article into its structure.

For replicability vs reproducibility, this guide follows the National Academies convention. Reproducibility asks whether the same data and code give results that agree. Replicability asks whether studies with their own data reach findings that agree on the same question.

Read the methods before relying on either label. Some fields use these words differently. State what data were used, what changed and what the researchers found in your note.

Atlas

Check which evidence was reused

Compare supplied study reports and keep data, method and result claims linked.

Define the terms before comparing reports

Start with data and analysis

The National Academies report draws its line between a check on existing data and a study with its own data. New data can test whether a finding persists beyond the first sample. A rerun checks what the steps yield from that sample.

Keep these two questions apart when you read a paper and its follow-up report. The table uses the terms in this guide. Other fields may name the checks in a different way.

CheckData usedMain question
ReproducibilityOriginal input dataDo the original steps yield consistent results?
ReplicabilityEach study's own dataDo findings agree on the same question?
Robustness in The Turing WaySame data; changed analysisDoes the finding hold under another analysis?

Table 1: The first two rows follow the National Academies; the third names a separate check in The Turing Way.

The Turing Way sets a changed analysis apart from a new-data study. For example, fitting a new model to the old data does not create a new sample. Record the data you reused and the model you changed.

Read the source's own definitions

Hans E Plesser's 2018 terminology discussion compares several earlier conventions. The original table below pairs labels with kinds of checks. The Claerbout and older ACM columns swap the two words.

Plesser's original 2018 table comparing Goodman, Claerbout and historical ACM terminology; Frontiers in Neuroinformatics.

Read across the rows: the labels do not line up in the same way in each scheme. This is evidence of the terminology problem as discussed in 2018. The ACM column is historical; do not use it as a statement of current ACM policy.

Hans E Plesser, Table 1, Reproducibility vs. Replicability, CC BY 4.0. Original PDF table cropped and converted to WebP.

Lorena Barba's review of the terms also shows how their use can conflict. Keep a source's meaning in your note when it differs from yours. State the action first, then explain how you mapped its label.

Measurement science uses a different scheme. Repeatability concerns precision under fixed conditions; measurement reproducibility concerns precision when conditions vary. That is a check on how close measured values are, rather than this guide's same-data versus new-data split.

Follow the data through four cases

Build the comparison from passages

Imagine a fictional classroom study of a short reading exercise. The first paper reports a mean improvement of 1.8 points. You have the paper, its methods file and a few follow-up reports. The numbers below show how to name the check. No study or code run was done for this guide.

For each report, find where the data came from, how they were used and what the check found. Then choose a label. Keep the question open when a passage is missing.

Follow-up reportData and stepsClassification hereWhat remains open
Rerun reports 1.8 pointsOriginal data and scriptReported computational reproductionWhich software versions and target outputs?
New class reports 1.6 pointsNew students; same exerciseReplication attemptAre estimates consistent given uncertainty?
Changed model reports 1.7 pointsOriginal data; new modelRobustness checkWhy was the model changed?
Data and script are sharedNo rerun result reportedMaterials availableHas anyone obtained the target result?

Table 2: These fictional rows track data, methods and results. The estimates alone do not establish success.

Correct the missing-data inference

The new class is a replication attempt because it supplies new data for the same question. A difference between 1.8 and 1.6 points does not by itself tell you whether the finding replicated. You need to know the study's design, how much doubt surrounds each estimate and its rule for judging a match.

Likewise, two rounded estimates of 1.8 could mask meaningful differences. Check what was measured and how the result was scored. A shared number does not mean that two studies tested the same claim.

The new-model row asks how a finding depends on the chosen analysis. It does not show that the effect held in a new class. Read a replication study as a test of a specified finding rather than any later paper with similar words.

The last row leaves a gap in the note. A data link and script let someone try the check. They do not report that it passed. Keep that gap even when the paper calls its work reproducible.

Decide what each check can establish

Separate access, attempt and outcome

Note whether you can access the files, whether someone tried the check and what they found. A link may exist while a file is missing or the code cannot run. A full run may still differ from the target output.

State which output was checked: one table, one figure or all the results. A matching figure does not show that every result was recovered. Keep any errors or manual changes beside the outputs that matched.

For a new-data study, read the question, sample and key conditions before judging what happened. If the follow-up changed both the people and task, explain those changes. Do not hide them under a simple “failed” label.

Keep conclusions within the evidence

A successful rerun supports a narrower claim than “the theory is true.” It shows what the stated data and steps produce. It cannot remove bias from the first sample or prove that the effect holds elsewhere.

A failed check needs a closer look too. Missing files, software differences, unclear methods or a different setting can affect the outcome. The National Academies discussion of replication cautions against judging a finding from one study alone.

Use the source-checking steps to keep reported claims distinct from your own checks. If a result is only described by the authors, say “the report states” so readers know that you did not run the check yourself.

Check post-publication critiques against the paper and any reply before treating them as evidence of failure. A comment can raise a useful question without settling the result.

Keep the cross-paper synthesis focused on the full evidence. One check that passed or failed cannot settle all claims about a finding. State the claim each check changes and what still needs testing.

Save the evidence comparison in Atlas

Create a project and add the original paper, method or data description and supplied follow-up reports. Wait for the sources to process. Give each report a clear name so you can distinguish the original work from a rerun or new-data study.

Open a chat and use @ to select those sources. Ask: “For each follow-up, cite the passages describing data origin, reused or changed methods and reported results. Under the same-data versus new-data convention, classify the check. Mark missing passages unknown and separate shared materials from a reported successful run.”

Open each citation and check the text around it. Correct “the study was reproduced” if the source only says that its code is available. Replace it with “materials shared; no successful rerun established by these sources.” Atlas helps compare descriptions here; it does not execute the analysis.

Choose New → Note to save the corrected comparison, source links and open questions. Check the Saved state, then revisit the note when another report arrives. The note helps you explain your judgment. It does not prove that the research is sound.

Atlas

Check which evidence was reused

Compare supplied study reports and keep data, method and result claims linked.

Frequently Asked Questions

Under the National Academies convention, reproducibility concerns consistent results from the same data and computational steps. Replicability concerns consistent findings across studies that each obtain their own data. Other sources may define the words differently.