Counterbalancing means varying the order in which people do a study's conditions. With two conditions, A and B, one group does A then B while another does B then A. This helps keep the comparison from depending on which task always comes first.
When a paper says its conditions were counterbalanced, look for the actual orders. Find how many people followed each order and how they were assigned to it. Those details tell you more than the label alone.
A study can balance which task comes first while leaving some pairs of tasks unequal.
Check reported orders in Atlas
Compare methods passages and save a checked condition-order table.
What counterbalancing changes
In a repeated measures study, the same person does several conditions. They may learn, grow tired, or approach a later task differently because of an earlier one.
If everyone does A first and B second, the effect of order is mixed with the difference between A and B. The experimental design textbook explains this problem through carryover effects: an earlier task can affect a later one.
Counterbalancing spreads the orders across people. Random assignment decides who enters each group. Record both steps: random assignment alone does not ensure equal group sizes. Cluster sampling instead picks groups to enter a sample; it does not assign people to task orders.
Imagine quiet and background-noise tasks done in opposite orders by two groups. These are teaching conditions, rather than a real study. The order-effects guide explains why earlier tasks can affect later scores.
Compare complete and partial arrangements
Complete counterbalancing uses every possible order. Two distinct conditions give AB and BA. Three give ABC, ACB, BAC, BCA, CAB and CBA. Four give 24 orders. More conditions soon mean many more orders to include.
Partial counterbalancing uses some of the possible orders. A Latin square is one approach. Each condition appears once in each row and once in each position across the square.
For three conditions, ABC, BCA and CAB meet that position rule. Each letter comes first, second and third once. These are three of the six possible orders, so the set is partial.
The MIT controlled-experiments lecture explains why an ordinary Latin square can leave pairwise learning unequal. Which task came just before another can matter too.
Count appearances in first, second and third place to check position balance. Then count which task comes just before each other task. These are different checks; a square can pass one without passing the other.
In the methods section, look for what the authors meant to balance. Use their listed orders or an appendix to work it out. A named design needs enough detail to show which orders the study used.
If the paper gives only the word counterbalanced, keep that limit clear. Write that the authors used the label, but did not give the exact orders in the material you read.
Read an AB/BA evidence table
Suppose a fictional study reports two groups of 12 people. One does quiet then noise; the other does noise then quiet. Its supplement lists the same orders but does not say how people entered the groups.
The table below keeps the reported orders apart from facts that the source does not supply. All passage labels are teaching examples.
| Group or item | Reported evidence | Order interpretation | Remaining question |
|---|---|---|---|
| Group 1 | Methods paragraph 2; 12 participants | Quiet then noise, equivalent to AB | Were these planned or analyzed counts? |
| Group 2 | Methods paragraph 2; 12 participants | Noise then quiet, equivalent to BA | Did every participant complete both tasks? |
| Sequence confirmation | Supplement section 1 repeats AB/BA | Both orders are represented | Were any deviations recorded? |
| Allocation procedure | No procedure in these passages | Allocation cannot be reconstructed | Was assignment random, alternating, or something else? |
Table 1: The fictional labels show what to record; they do not refer to a published study.
With these equal group counts, quiet comes first for 12 people and second for 12. Noise does too. That supports position balance in this teaching example.
It does not show how people were assigned, whether anyone was left out of the analysis, or whether one task had lasting effects.
Suppose the supplement says three people did not finish the second task. Keep the planned counts and add that fact. The plan alone no longer tells you which tasks each person actually did.
To read the final comparison, you need to know which group lost people and which scores stayed in the analysis.
A claim such as “the study removed order bias” goes beyond these passages. A supported account is: “The authors report equal AB and BA groups. The passages do not say how people were assigned.” That keeps the known facts and the missing method together.
Check what remains unbalanced
In the cyclic example ABC, BCA and CAB, A comes just before B in ABC and CAB. B comes just before A in none of the three orders. Yet each letter still appears once in each position.
MacKenzie's research note discusses balanced Latin squares that address unequal pairs. The exact orders matter when you read what kind of balance a paper claims.
Ask whether the design addresses the task pairs that could affect the next score. The words Latin square do not let you infer a specific set of rows when those rows are missing.
Group sizes matter too. A square may balance positions across its rows. If many more people follow one row, the positions are no longer used equally often across people.
Order labels do not settle causal assumptions. A 2025 methodological preprint studies carryover and identification assumptions in these designs: what must hold to infer a condition's effect from the comparison.
The preprint supports caution about treating counterbalancing as a guarantee. It does not give a universal rule to add a particular test or gap between tasks.
Keep three facts in your review notes: the orders used, the balance you can check, and the effects or assumptions the authors studied.
If a missing supplement stops you checking the balance, keep that gap. Do not fill it with a textbook example. If you also need to know whether the study changed what it meant to change, a manipulation check addresses that separate question.
Reconstruct reported orders in Atlas
Add the methods, supplement and protocol to an Atlas project, where you have permission to use them. Wait for the sources to finish processing. Open a chat and use @ to select the methods and supplement.
Ask: “List each group, task order, count and reason for counterbalancing. Cite each entry. Flag missing assignment or completion details, and keep reported facts apart from your interpretation.”
Open each citation and read the text around it. Check the source name and task order. Also check whether a count refers to people recruited, people assigned, or people whose scores were analyzed.
If a citation opens the document without showing the exact passage, find the section yourself before accepting the entry.

Atlas interface screenshot showing an open citation preview. The visible ColPali paper by Manuel Faysse and colleagues is CC0 1.0; it is unrelated to the task-order example. Lossless WebP re-encoding left the content unchanged.
An answer might call the fictional AB/BA groups randomly assigned, though the passages give only orders and sizes. Replace that claim with “assignment method not reported in the checked passages.” Ask a follow-up for the exact wording, or compare the protocol's plan with the final report.
Use New then Note to save the corrected table, source references and gaps. Wait for Saved before closing the note.
Add later supplements to the same note. The research synthesis workflow helps when you compare papers that use different orders.
Write a bounded methods interpretation
A useful note names the conditions, orders, counts and source passages together. In the teaching example, both groups have 12 people and use opposite quiet/noise orders.
Those counts show equal use of each position in the example. How people entered the groups and whether they finished remain open questions. Keep those gaps beside the claim they limit.
If you are judging the research, rather than describing it, state what further evidence you need for that judgment.
Questions about lasting learning or fatigue need the study's design rationale and analysis. Opposite orders show an arrangement. They do not prove that every relevant effect of task order was removed.
Before reusing the note, check the passage for each study-specific claim. Keep missing details visible and your judgment apart from reported facts. New evidence can then change the judgment without erasing what the authors first reported.
Check reported orders in Atlas
Compare methods passages and save a checked condition-order table.

