Parallel forms reliability asks how well two test versions measure the same trait. The forms use different questions but aim to serve the same purpose. The claim needs both a plan for matching content and results that show how the forms behave.
A shared title or question count leaves much unknown. Read the test guides, study schedule and results together. Atlas can help compare those sources and save a note that keeps the intended match separate from what the study found.
Check the evidence across test forms
Compare form descriptions and results, then save the supported similarities.
What parallel forms reliability means
Two forms may test the same domain with new questions. For example, a word test could use two word sets to assess the same skill. That goal does not prove that the forms give scores you can swap.
The SAGE entry distinguishes alternate forms in everyday use from strict parallel tests in classical test theory. Strict parallel tests have matching latent true scores and error variances, with independent errors.
These technical terms describe the assumed true-score and error structure. They go beyond matching names, lengths or topics. Keep them exact when an author makes that strict claim; do not apply the label just because both forms have 20 questions.
In practice, authors may mean that they built forms to approach a match. Record what they checked before you adopt the stronger claim. The Scribbr overview helps separate this design from other repeat checks.
Giving one unchanged test twice concerns test retest reliability. Looking at links among items in one test concerns internal consistency. Changing the items, test version or testing date changes the question the study answers.
Compare the content before the coefficient
Start with the trait and content plan. Which skills does each form cover? How much weight does each skill get? Read the answer format, score range and scoring rules too. These details show what the authors tried to match.
The same question count is easy to check. The demands of the questions need closer reading. One form might ask people to recall facts while the other asks them to solve a problem in several steps. Both may cover the same broad topic without posing the same task.
Ketterlin-Geller and colleagues describe how they built parallel math items. Their Grade 5 blueprint sets out how many items to assign to each skill before cloning questions. It covers multiplication, division, decimal operations and other parts of the domain.

Figure 3 is reproduced unchanged from Ketterlin-Geller, Sparks and McMurrer (2022) under CC BY 4.0.
The blueprint links each skill to an item count. It plans content coverage, but matching difficulty still needs empirical tests. Keep the plan and observed scores separate in your note.
Also check whether items overlap. The same questions in a new order still have the same content.
The SAGE discussion warns against treating that change as distinct item sets. Statology also notes that random halves can differ in difficulty. A way of making forms is not proof of the result.
Read the administration schedule
Find who took each form and when. If people took the forms on different days, both time and content may affect the scores. Learning, recall or tiredness may change how people do on the second form.
Order matters too. If everyone takes A and then B, form and occasion line up. Giving some people B first can help address an order effect, but you need to read how the study did it. A claim that order was balanced should have a clear methods basis.
The assessment chapter links alternate-form checks to the time gap and carryover. Use it to frame the question, then find the reason for this study's schedule. A gap suited to one trait may be a poor fit for another.
Keep the exact statistic, score scale and uncertainty with the result. A correlation tells you how scores move together. It does not, on its own, show that each person gets the same score on both forms.
Imagine made-up A scores of 20, 30 and 40, with B scores of 30, 40 and 50. Their ranking agrees perfectly, yet every B score is ten points higher.
The arithmetic example shows why the same rank need not mean the same score. A decision to swap scores needs evidence for that specific use.
A worked form comparison note
Imagine two invented reading tests with 20 questions and the same score range. Form A focuses on fact recall, while B has more questions that ask people to infer meaning. The report says students took A and then B a week later. It gives a correlation but no check of order effects or score equating.
The table keeps the plan and result apart. None of its details represent a real study or computed result. It is a way to see which claims a source can support and which still need work.
| Evidence | Reported detail | Supported interpretation |
|---|---|---|
| Length and scale | Both have 20 items and matching score ranges | Some surface features match |
| Content | Different balance of recall and inference | The content match remains open |
| Schedule | A followed by B after a week | Both form and time may affect results |
| Comparison | Correlation without equating evidence | Linked scores need not be scores you can swap |
Table 1: For a real paper, add a page or section for each fact. If the paper leaves a detail out, keep it missing rather than borrow it from another test's guide. The ETS guide gives context on designs and error sources; the actual study must answer the question about its own forms.
Revise the note when a content map or new analysis appears. A fuller result may support a stronger claim. The note should still distinguish the planned match from the observed scores and the proposed use.
Check form evidence in Atlas
Add the permitted test guides, study plan and results paper to one Atlas project. Start a chat and use @ in Ask a question to select each source, then ask:
Compare these alternate forms. Find the trait, content coverage, item overlap, score scale, order, time gap and reported result. Cite each detail. Keep the planned match separate from what was observed. Mark missing facts as not reported. Do not calculate or equate scores.
Select Send and open the citations beside claims about content and timing. Read the text around them. If the answer says both forms share a plan, check that the source actually says so. Similar names are too weak a basis for that claim.
A high correlation may prompt the answer to call scores interchangeable. Ask which analysis supports swapping one score for the other. If none is shown, correct the claim and keep the gap in the note.
Select New and Note, then save the corrected note with source pages and open questions. Wait for Saved before closing. Atlas supports reading and comparison. Researchers build tests, choose analyses, equate scores and judge whether a score fits its intended use.
Keep the equivalence claim bounded
Name the forms, group, schedule and result when you write up the finding. The SAGE distinction helps avoid strict parallelism claims when a paper only shows an approximate match.
Read the evidence again for a new group or changed form. New wording, language or test conditions may limit how much an old result tells you. If one person scoring the same cases twice is the concern, use the separate intra-rater reliability workflow instead.
Check the evidence across test forms
Compare form descriptions and results, then save the supported similarities.

