Reliability vs validity asks two questions: do the scores hold up when you repeat a measure, and does the evidence support what you say they mean? A score can stay much the same while measuring the wrong thing.
Imagine a scale that gives the same reading each time but adds five kilograms. It repeats the result well, yet a known weight reveals the bias.
For a survey, the question of score meaning is broader. Traits such as confidence have no single known weight against which you can check them.
When reading a paper, keep its findings about repeatability apart from its claims about score meaning. The worked example shows how to record both and keep gaps in the evidence visible.
Keep consistency and validity claims separate
Check what your measurement papers support before saving the comparison.
What separates reliability from validity
Reliability concerns how well a measure repeats under the conditions being studied. Validity concerns support for what the scores mean and how they will be used. The Testing Standards put score meaning and use at the center of validity. Support for one use need not justify another.
Suppose a survey asks whether students feel sure they can write a review. It might support a claim about their confidence in that task.
Using the score to rank their research skills needs a broader case. Do the questions cover those skills? Does feeling sure mean that someone can do the task well?
The distinction changes the claim you write. “Responses were consistent across items” reports one property. “The score measures research skill” asserts a meaning that needs more support. Keep that second claim open while you look for the evidence.
Identify the consistency being tested
Before reading the number, find what the authors repeated or compared. Did they repeat the test, compare its items, or compare raters? These answer distinct questions. Keep the method name beside the result in your note.
Can reliability exist without validity?
Yes. A repeated result does not tell you which trait it measures. The BCcampus chapter gives the example of using finger length to infer self-esteem. A precise ruler cannot make that claim sound.
For a survey, the mismatch may be less clear. Several questions about writing confidence can give stable answers while leaving research skills out. The score may be useful for the narrow trait even when it cannot support the broad claim.
Consistency over time
Test-retest evidence compares scores from the same people at two times. Check the time gap and why the trait should stay stable. A changed score could reflect real change, error, or both.
Mood may change over a month. The same scores at both times would not always be the goal. For a stable trait over a short gap, change raises another question. Note what the study expected to stay the same.
Consistency across items
Internal consistency asks how responses to items within a test relate. Cronbach's alpha is one common estimate. It cannot by itself prove that the items measure one trait or cover all of that trait. The scale-development primer treats scale structure and reliability checks as distinct steps. Read both parts of the paper before judging the total score.
Asking much the same thing several times can leave out the rest of the trait. Find the items and the score they form. If the study did not repeat the test, its item-level result is not evidence of stability over time.
Consistency across raters
Inter-rater evidence compares judgments made by raters. Find the rating rules, what they judged, their training, and the measure of agreement. Two raters can agree while using a rule that does not fit the task. Agreement answers whether they judge alike. It does not settle whether the rule captures the trait you want to study.
Look for the interpretation evidence
Start with the claim the score should support. Does it describe a trait, compare groups, or track change? Ask what findings would support that claim and what else might explain them.
The COSMIN taxonomy sets out distinct properties and study needs for health measures. It names content, construct, and criterion validity.
The Testing Standards treat validity as a single argument built from several sources of evidence. Use the terms in your paper, but explain what each finding supports.
Coverage and score meaning
Content evidence asks whether the test covers the trait it claims to measure. A test of research skill needs a reason for the skills it includes. Questions about confidence in writing may cover a smaller part well.
Other evidence may show how people read the questions, how items form a scale, or how scores relate to other traits. What you need depends on the claim.
Read content validity versus construct validity when you need to distinguish test coverage from the broader case for score meaning.
Keep the context attached
Check the group, language, test version, and how the test was given. A new language, item set, or use may raise questions the first study did not address. Keep these details in your note.
To compare groups, read the paper's measurement invariance evidence. Scores that repeat well within each group may still have meanings that differ across groups.
Work through a measurement report
Consider a fictional study of a Research Confidence Questionnaire at one college. It asks about finding papers, reading methods, and writing a review.
The authors report internal consistency and call the score a measure of overall research skill. This teaching example supplies no real participant data or coefficient.
Read the claim against the evidence below. Each row identifies what can be retained and what still needs a supporting passage.
| Fictional report detail | Evidence question | What the note can retain | What remains unresolved |
|---|---|---|---|
| Items have a reported internal-consistency estimate | Are item responses consistent? | The report presents evidence about consistency across items | Scale structure and interpretation require separate support |
| Items ask about perceived confidence | What does the score represent? | The questions concern self-perceived confidence | They do not directly demonstrate research competence |
| No items concern practical analysis tasks | Does the content cover the claimed domain? | The item set leaves the proposed broader domain incompletely described | Find the authors' domain definition and coverage rationale |
| One cohort completed one version | Where does the evidence apply? | The reported evidence belongs to that cohort and version | Stability, other populations, and changed versions are not established here |
Table 1: This fictional report separates a consistency finding from a broader interpretation that still needs support.
The missing skill tasks do not prove that a confidence score is invalid. They show a mismatch with the broad claim in this example. The researcher could narrow what the score means to match the questions asked.
Repair the inference
The claim that the test is valid because its items were internally consistent goes beyond the report. A more careful note would record that the authors found consistency across items for this version in one group. The material reviewed does not establish that the score measures broader research skill.
The new sentence keeps the finding, context, and open question. If another paper supplies support, add it with its source location. A research-paper analysis workflow helps keep findings apart from claims made about them.
Build a checked note in Atlas
Create a project with the paper, its scoring guide, and any studies you want to compare. Wait until each source has finished processing. Add the full report and supplements when the abstract leaves out test details.
Start a chat and type @ to select the sources. Ask: “Split evidence about repeatability from evidence about score meaning. For each claim, give the test version, group, method, result, limit, and citation. Mark missing evidence.”
Open each numbered citation beside a claim. Read the source text around it to check which property was studied. If the answer says the score proves skill, ask for that supporting passage. Correct the claim when the source does not supply it.
The capture shows a real Atlas answer beside an open source. The paper is about AI scientific discovery. It shows citation checking, not a study of the fictional test.
For your own papers, use the open passage to check what was tested, what was found, and which limits apply.

Keep the source open while checking whether an answer extends beyond the reported evidence.
The visible paper is Yamada et al.'s AI Scientist-v2 study, licensed CC BY 4.0. Its page appears unchanged within the Atlas capture.
Choose New, then Note, and save the checked findings. Split the note into repeatability, score meaning, and open questions. Add the sources and locations you checked, then wait for Saved before closing it.
Atlas helps compare supplied sources and save notes. Researchers choose measures and judge the claims. Atlas does not calculate reliability or validate a scale.
When several studies add distinct evidence, a source-grounded literature review can keep each study's contribution visible. Use paper-reading tools to assist with reading, while retaining your final judgment.
Write the claim the evidence supports
Before using the note in a paper, check each claim against its source. Does it keep the test version and group in view? Does it distinguish the finding from your judgment? Name the support provided instead of using an unqualified “validated instrument” label.
Keep gaps visible. “Not reported here” differs from “tested and unsupported.” The first may need another source; the second calls for close reading of the method and result.
Keep that distinction when writing a research-paper summary. If the gap concerns links with related traits, the convergent validity guide explains which evidence to look for.
Revisit the note when you change the test, group, language, or use. A reviewer should be able to trace which finding you kept, which meaning it supports, and which question still needs study.
Keep consistency and validity claims separate
Check what your measurement papers support before saving the comparison.

