Skip to main content

Blog

Content Validity vs Construct Validity: Read the Claims

Compare content validity vs construct validity with a worked domain-coverage table. Check item content, score meaning, and gaps in a paper's evidence.

Semantic Map: Visualize the topic from new angles.
Knowledge Map: Deconstruct the article into its structure.

Content validity vs construct validity separates what a test covers from the case for what its scores mean. Content evidence asks whether the questions or tasks represent the concept. Construct evidence asks whether the score behaves as the theory would lead you to expect.

The two questions overlap. A test cannot make a strong claim about a broad skill if it leaves out an important part of that skill. Yet complete-looking content does not by itself settle how people read the questions or what the total score means.

Use the worked example to compare those evidence claims in a paper. The aim is a source-linked table you can inspect and revise, with gaps kept visible.

Atlas

Check which validity claim the paper supports

Compare the item coverage with the evidence for score meaning.

What each validity question asks

Content validity concerns the match between test content and its intended domain. A domain is the set of aspects the test aims to cover. For a test of planning confidence, that might include setting goals, allocating time, and revising a plan.

Construct validity concerns support for the concept assigned to scores. A construct is the idea you want to measure, such as confidence. You need a reason to think the score reflects that idea rather than another trait or a feature of the test.

The Testing Standards treat validity as an argument about score meaning and use, drawing on several sources of evidence. Content contributes to that argument. It is not a rival option that replaces the rest.

The COSMIN framework names content and construct validity as distinct properties of health measures. When reading a paper, state which framework it uses. The label helps you track what was checked; it does not settle the quality of the evidence.

Read the domain and item coverage

Find the authors' definition before reading the items. If the domain is unclear, you cannot tell what is missing or outside its scope. A familiar scale name is not a substitute for a definition.

Check what belongs in the test

Map each question or task to the stated domain. Does it ask about a relevant aspect? Does the item set omit one? The BCcampus measurement chapter uses concept definitions to explain this coverage question.

For the planning example, several questions about setting goals do not cover time allocation just because they use similar language. More items need not mean broader coverage. Record the aspect each one addresses rather than counting the questions alone.

Read how the review was done

For scores based on patients' reports, the original COSMIN content study asks whether items fit the trait, cover it, and make sense to the people who answer them. These checks concern a given group and use.

Find who reviewed the test, what they were asked, and what they found. A panel may think an item fits while a reader takes it to mean something else. Keep both views and any changes made after review.

Keep the version in view

If the team removes items after the review, check whether the final set still covers the domain. Evidence for the long version does not automatically describe the short version. Add both version names to your comparison so that a later draft does not merge them.

A questionnaire-design workflow can help with item wording and development. Here, the job is narrower: read what the completed study reports, then mark which coverage claims it supports.

Check evidence for score meaning

Look for claims the authors made about the score before interpreting the result. What patterns should you expect if the score measures the stated concept? What patterns would be surprising? The reasons belong beside the findings in your note.

Read the pattern among items

A factor model studies how responses to included items relate. It can inform the case for a scale or subscales. The scale-development primer treats this as part of scale construction and evaluation.

A model may fit the included items while an aspect of the domain is absent. If all the planning questions concern goals, a model of those responses cannot show that the test covers time management too. Check the content definition as well as the model.

Links with related scores can add support when theory predicts them. Convergent validity asks about those links. Two scores with much the same name need not measure the same thing.

The COSMIN reporting guide asks authors to state which way the link should go, how strong it should be, and why. Its scope is patient-reported scores. Use that lesson with its scope in view.

Distinguish repetition from meaning

A score may repeat well while leaving out part of the domain or capturing the wrong trait. Reliability versus validity shows why that matters. Keep a favorable finding, but name what the study tested.

To compare groups, read the paper's measurement invariance evidence. Does the test work in the same way for each group? This is a distinct question from what the items cover.

Compare two fictional study claims

Imagine a fictional Planning Confidence Scale. It should cover choosing goals, making time, and revising a plan after feedback. One report reviews the items; a second looks at how scores relate. This example uses no real scores, people, or fitted models.

Use this table to compare the claim each passage can support and the next source check it requires.

Fictional report passageEvidence familySupported readingNext check
Experts found the goal and time items relevantContent relevanceIncluded items fit two stated aspectsWere feedback and revision covered?
No final item asks about revising a planDomain coverageThe final item set leaves one stated aspect absentDid the authors narrow the domain or revise the set?
A factor model fits the included itemsInternal structureThe report supports a pattern among those responsesDoes that pattern justify the claimed total score?
Scores relate to an existing planning measureRelationships with other variablesA reported link may fit the theoryCheck the definitions, predicted link, and method overlap

Table 1: The fictional findings support distinct claims; no single row establishes the whole scale's validity.

Repair the combined claim

An overbroad note says: “Experts approved the items and the model fit, so the scale covers all planning confidence.” The missing feedback aspect remains a gap. Model fit does not make it part of the test.

A more careful note says: “The report supports the goal and time items and shows how their scores relate. The final set leaves out plan revision as defined here. The broad score claim still needs that gap resolved.”

Keep the task distinct from certification

The table helps you read and compare reports. It does not replace a full scale review or prove which test to choose. You may need more studies and help with the methods before making that choice.

When using several reports, a source-grounded literature review should preserve which study supplies content evidence and which supplies score evidence. Do not turn their combined presence into stronger support than their methods allow.

Check the comparison in Atlas

Create a project with the domain definition, final item set, development report, and validation paper. Wait for the sources to finish processing. The final item set matters: a draft version can differ from the version used to produce the scores.

Start a chat and use @ to select those sources. Ask: “Compare domain coverage with the evidence for score meaning. Give one row per claim, its source passage, test version, group, finding, and limit. Mark missing evidence rather than filling it in.”

Open each numbered citation. If a response uses factor fit as proof that no content is missing, compare the claim with the item list and domain definition. Ask for a narrower answer and correct the wording in your working table.

The screenshot shows a real Atlas answer beside an open AI-science paper. It illustrates the source-inspection step, not content ratings or factor-analysis output for the fictional scale. Read the cited passage before treating any row as checked.

Atlas answer beside an open research paper, illustrating how to inspect a cited claim against its source

Inspect the source context before merging claims from a development report and a validation paper.

The visible paper is Yamada et al.'s AI Scientist-v2 study, licensed CC BY 4.0. Its page appears unchanged within the Atlas capture.

Choose New, then Note, and save the corrected table. Keep the source locations, version names, and gaps with it. Wait for Saved before closing the note; a useful title might name both the scale and the question left open.

Atlas can help compare supplied text and save your checked notes. Choose the content-index calculation and check its inputs and results. Specify and evaluate any factor model, then judge whether the full evidence supports the instrument's intended use. A paper-analysis workflow helps you retain that distinction between reported findings and your own judgment.

Preserve the evidence gaps

Before using the table in a draft, check which evidence is missing and which points to a problem. “No content review is reported” differs from “the review found a missing domain.” These call for distinct next steps.

For missing detail, check supplements or a new source. For a problem found in a study, read the method, version, and authors' response. Keep that finding in the paper summary; do not erase it with a broad validity label.

Revisit the table if the authors change items, redefine the domain, or use the score with a new group. Coverage and score meaning remain connected, but your note should show which source supports each claim.

Atlas

Check which validity claim the paper supports

Compare the item coverage with the evidence for score meaning.

Frequently Asked Questions

Content validity concerns whether test content represents the concept it should measure. Construct validity concerns the case for the meaning of its scores. Content evidence can contribute to that broader case.