A replication study uses new data to test a claim from earlier research. Its value comes from what the new evidence tells you about that claim, including when it challenges the earlier account.
When you read an original paper beside its replication, start with the claim and the features of each test. Then compare what each paper reports, what changed and what remains unclear.
The fictional example below shows why a changed test delay matters more than a tidy “same result” label. Atlas can help build the source-linked comparison; you check the rows and judge their meaning.
Check an original and its replication
Compare the claim, method and result without losing the departures.
What a replication study tests
A replication should tell you something about an earlier claim. Nosek and Errington propose that a new test should work both ways.
A result that fits the claim should strengthen your trust in it; a result that does not should weaken it. This is their proposed definition, and fields vary in how they name these tests.
Repeating a task is not enough if the new study asks a different question. State which earlier result is at stake and why the new evidence bears on it. A claim about recall after a week, for example, needs a test that can inform recall at that delay.
Here, replication means testing with new data. If you rerun the same data and code to check the result, that is computational reproducibility.
Define which task you mean, since other fields may use the words differently. Reading two papers lets you compare their research; conducting a new replication would require a new study.
Fix the claim before comparing papers
Locate the claim in the original paper and keep its scope. Who was studied, what was compared, which outcome was measured and when? An abstract may compress those details.
Read the methods and results, along with the protocol and supplements when available. Verywell Mind's guide starts the replication workflow with the original hypothesis, participants and methods.
In the fictional example below, paper O claims that a practice task improves recall after seven days compared with rereading. The claim concerns a stated delay and a stated comparison. It does not say every form of practice improves all learning outcomes.
Check whether the claim was planned or developed after seeing the data. Mark that status rather than treating every result as the main hypothesis. If you are planning a new test, turn the target into research objectives and state which outcomes would inform it before seeing results.
A broader theory can help explain why a detail matters. A theoretical framework should connect the claim to that rationale; it cannot make a mismatched measure equivalent by itself. Keep the theory, the tested claim and the result distinct in your note.
For a claim about why a pattern occurs, explanatory research helps you state the proposed reason, rival account and evidence still needed. Specify which part of that explanation the new test can inform.
Check the study features that matter
“Same method” can hide several differences. Compare who or what was studied, the task or treatment, the outcome measure and the setting. Nosek and Errington's Figure 1 uses those four features to describe the conditions under which a claim is tested.
The original key below names these features. It gives you a compact way to inspect the papers before interpreting their results.

Nosek and Errington, Figure 1, PDF page 5, name four test features. This excerpt is from their real methods perspective, separate from the fictional papers used below.
For each feature, record what each paper says and where it says it. “Not reported” is useful when a needed detail is absent. Do not fill that gap from what similar studies usually do. If you consult a supplement, keep its version and locator too.
If either paper compares groups formed before the study, use the causal comparative research guide to check how those groups formed. Keep selection and other possible explanations beside the result before treating a repeated group difference as a treatment effect.
Labels such as direct, close and conceptual replication describe broad aims, but they do not replace this comparison. A direct or close attempt usually retains key features; a conceptual attempt changes how the claim is tested.
McManus's field methods commentary explains why the actual similarities and changes need clear reporting. Its detailed change categories belong to its field, so avoid imposing them as universal counts.
A worked original-replication comparison
Papers O and R, their locators and their statements are fictional teaching materials about practice and recall. No findings here are drawn from a real learning study. In this packet, O is the original paper and R describes itself as a replication with new data.
Use the located details below to check which parts match and which change the claim being tested.
| Feature | Original paper O | Replication paper R | Consequence for the comparison |
|---|---|---|---|
| Group | Page 2: students in one course | Page 2: working adults recruited online | The group changed; judge its relevance to the claim |
| Task and setting | Page 3: ten-minute practice in a lab | Page 3: ten-minute practice at home | The duration matches; supervision and setting differ |
| Outcome timing | Page 4: the stated recall quiz after seven days | Page 4: the stated quiz after one day | The new delay does not directly test seven-day recall |
| Result summary | Page 5: a significant advantage is reported | Page 5: a nonsignificant advantage is reported | Obtain effect estimates and uncertainty before a verdict |
Table 1: The fictional table separates source details from their implications. A matching task name or duration cannot settle whether the papers test the same claim.
The third row should change your conclusion. R may inform short-delay recall, but its one-day test does not directly answer O's seven-day claim. Record that limit even if both papers used the same quiz. A measure includes when it is taken as well as what it asks.
The last row leaves a second question open. You know the authors' significance labels, but those labels do not show the size or precision of either effect. Do not write “the effect disappeared” from that summary alone. The packet needs fuller results before you can assess that part of the comparison.
Read results beyond significance labels
Check effect size and uncertainty
Record how large an effect each paper reports and how precise that estimate is. Keep the scale, group sizes and method with it so you know what the numbers mean.
A confidence interval is one way authors express uncertainty under their model. If they report one, copy its limits and check which effect it describes.
Interpret thresholds with care
The labels “significant” and “nonsignificant” say whether results met a threshold under a test. They do not tell you that the two study effects differ.
Group sizes and the precision of the estimates can change those labels. To assess how the effects relate, you need a sound statistical method that fits the papers and their data.
Read the report's assessment
The APS Registered Replication Reports model uses a shared plan across teams. It asks for each team's results and an effect estimate across the studies, moving away from a simple success/failure label.
This model has a shared protocol; you still need to judge whether it makes sense to combine any other set of papers.
For the fictional packet, seek the full results and supplements from O and R before judging how their effects compare. If you cannot get them, save that gap in the note. The word “nonsignificant” does not supply a missing effect size, however closely you read it.
Judge departures without excusing every outcome
A departure is a change you can document between the studies. Check its role before judging it.
The authors may have planned a new group to test how far a claim extends, or changed a delay to ask about a new outcome. If the method is unclear, mark what you cannot tell from the report.
Explain why a change bears on the claim. In O and R, delay matters because the claim names seven-day recall. The new adult group may matter too, but give a reason to expect it would change the effect. The fact that the people differ does not, on its own, explain a result you did not expect.
Apply the same standard to results you like. Do not accept a supportive outcome as decisive, then dismiss a challenging outcome solely because the setting differs. The original claim-testing perspective makes both directions part of the test.
For a planned study, record intended changes and their reasons before you collect data. When reading papers, say when you learned about a change and what lets you judge it. McManus's reporting discussion asks authors to make those reasons and changes clear so readers can assess them.
Compare the supplied papers in Atlas
Prepare both papers
Add the original paper, replication paper and any permitted protocols or supplements to one Atlas project. Use Add a source and Upload files, wait for processing readiness and check that the needed pages are readable.
Give each item a clear name so a citation to a supplement cannot be mistaken for the main paper.
Inspect and save the comparison
- Mention the selected sources with
@. Ask for a table comparing the claim, group, task, setting, outcome timing and reported result, with a supporting citation for each side. - Open each numbered citation and read the surrounding passage. Check that the row points to the right paper and distinguishes a planned method from what was actually done. Exact source positions depend on available information.
- Correct the comparison. Keep missing details, changed measures and the authors' result labels visible. Ask Atlas to separate source facts from interpretations that go beyond them.
- Create a note through New and Note. Save the checked table, claim, source versions and unresolved questions, then confirm Saved after editing.
For O and R, ask which passage supports calling the test delay the same. If the answer treats the quiz name as enough, correct the row with both page-4 locators.
Atlas helps compare supplied text; you retain the design and statistical judgment. A clear system for organizing research notes helps you return to those checks later.
Save a conclusion with its limits
You can save a bounded note for the fictional packet: “R reports new data on the practice task with a new group and setting. Its one-day quiz does not directly test O's seven-day claim. We still need the full results to judge how the effects compare.” Keep the table and missing details beside that note.
This gives a later reader the reasons behind your view: what was compared, where the papers differ and what could change your judgment. When you discuss research implications, keep the claim at that strength. One pair of papers can inform a theory, but it cannot settle every claim the theory makes.
Check an original and its replication
Compare the claim, method and result without losing the departures.

