Face validity asks whether a measure appears to test the concept it claims to measure. A question can look relevant without evidence that its scores work as intended. Keep that gap visible when reading an instrument description.
Your task is to compare the stated concept, the actual wording, and the reported judgment about its fit. The result should be a short caution note, not a verdict that the measure has been validated.
Check the claim against the item
Compare item wording, the stated concept, and the reported review.
What face validity tells you
An item is a question or task within a measure. A construct is the concept the measure aims to capture, such as confidence when reading unfamiliar text. Face validity concerns whether the item seems to fit that concept on inspection.
The original textbook treats this fit as weak evidence. An item that looks right is a starting point to check, rather than proof of what a score means. A clear question can still yield a score that fails to capture the intended concept.
The judgment can be informal or summarized through ratings. Numbers can show how reviewers judged relevance, but they do not turn appearance into proof that the measure works. The textbook explicitly allows ratings of how well a measure appears to fit its intended concept.
Experts and intended users may hold different views. Scribbr's review example shows why a question that seems clear to the person who wrote it may confuse those asked to answer it.
Read the item and review together
Read the measure's text with its stated purpose and review report. You can then say what was judged without giving the claim more force than the source supports. Keep the exact question at hand so a label does not replace its wording.
| Check | Evidence to locate | Limit to retain |
|---|---|---|
| Concept | What the authors intend the score to mean | A broad concept may need more than one item |
| Item | Exact wording, task, and response options | A relevant topic can still be unclear or incomplete |
| Review | Who judged it, what they were asked, and which version they saw | A judgment on one version need not apply to a changed version |
Table 1: Three checks for an apparent-fit note; they do not replace an instrument validation study.
The expert-review methods paper gives judges context about what the items aim to measure. Look for that context in your paper. A bare “face validity confirmed” claim does not tell you what the review involved.
Keep who reviewed the items and what they judged in the same note. Saying that an item is clear does not mean it covers the whole concept. The author's own view also differs from a review by the people who would use the measure.
Simply Psychology's definition starts with whether the item looks relevant. To judge how scores work, read the study evidence. The look of an item cannot settle what its scores mean.
Worked example with a plausible item
Imagine a fictional scale about confidence when reading a new research text. One item says, “I feel able to explain the main claim after reading a new article.” The scale and review below are made up for teaching; no study was run.
The wording seems to fit reading confidence because it asks about a person's sense of ability. The color of the paper or a choice of desk would have a less clear link. You can recognize this fit without claiming that the scale works well.
Suppose the report says intended users found the item relevant and clear. It then claims, “The scale was validated for measuring reading confidence.” The first statement describes what users judged. The second needs more evidence about how the scale works.
A corrected note might read: “The report says users found this item relevant and clear for reading confidence. This supports the fit of the wording they saw. It does not show that the whole concept is covered or how the scores relate to other evidence.”
The expert-review paper keeps looking relevant apart from covering the full concept. Use that distinction to bound your note. Do not add ratings, counts, or test results absent from the report.
If the item changes to “I always understand every article,” revisit the note. The words “always” and “every” change what people are asked to judge. A review of the earlier sentence cannot simply apply to this new version.
Separate apparent fit from stronger evidence
Content validity asks whether the measure covers the intended domain: all the parts the concept is meant to include. The textbook's coverage explanation gives this a broader task than asking whether one item looks relevant.
In the fictional scale, a main-claim item may leave out other parts of what the authors intend to measure. Check their definition to see which parts matter. Do not broaden “reading confidence” into every reading skill you can think of.
Reliability asks whether a measure is consistent. Steady answers need not support the intended meaning of the scores. The textbook keeps those questions apart: being consistent does not by itself show that a measure works as claimed.
You also need evidence about how scores relate to other things the theory predicts. The guide to convergent validity addresses one such link. A face judgment alone cannot show that the scores have this link.
Poor fit on the surface does not prove that a measure cannot work. Some well-supported measures have less obvious links to the concept. Read the wider evidence before dismissing a measure just because its items look indirect.
For unfamiliar wording or a new user group, Scribbr discusses renewed review. Treat this as a reason to check the context, rather than a claim that one favorable review settles all future uses.
Compare the item sources in Atlas
Add the items, concept description, and review account to one project. Use only sources you have the right to share. Keep the version and answer choices so you compare the wording that people actually saw.
The research paper analysis workflow helps connect these passages with the paper's broader claims.
- Select the sources with @. Ask: “What concept should this item measure? What is its wording, who reviewed it, and what did they judge? Keep the fit of the item separate from evidence about how scores work.”
- Open every citation and read the surrounding text. Check whether the cited passage concerns the item, the review, or a separate validation study.
- If the answer calls the scale “validated” from a face judgment alone, ask for the study evidence. Replace the claim with what the passage says people judged.
- Create a note with New → Note. Save the concept, item version, reported judgment, correction, and missing evidence with citations. Wait for Saved.
The capture below shows a cited answer beside an unrelated paper. It shows how to open source context, rather than our fictional scale or a test of its scores. Read the passage to check the strength of the answer's claim.

Read what the passage actually establishes before describing an item review as instrument validation.
Screenshot: Atlas. Embedded paper: Yamada et al., The AI Scientist-v2, CC BY 4.0. Paper content unchanged within the product capture.
Save a precise validity caution
Save the wording, what the users or experts judged, and the limit on that claim. If the paper omits the review task or version, leave the gap open. A clear account of what is missing helps you decide which source to read next.
When you synthesize several papers, keep the forms of evidence apart. An item that looks the same in another study does not bring its review or score evidence into this one. The wording, users, and task may differ.
Within your literature review process, use the caution note to qualify claims about measures. Stronger evidence can revise the note later; apparent relevance alone should not close the question.
Check the claim against the item
Compare item wording, the stated concept, and the reported review.

