Skip to main content

Blog

Intra-Rater Reliability Through Repeated Judgment Evidence

Intra-rater reliability checks one rater's repeated judgments. Trace what was repeated, the rating interval, blinding and reported agreement in a study.

Semantic Map: Visualize the topic from new angles.
Knowledge Map: Deconstruct the article into its structure.

Intra-rater reliability asks how well the same rater judges the same targets again. Read what was done twice before you apply the result. Scoring a saved image twice and taking a new image before scoring it test different parts of the process.

A cited note can keep that distinction clear. Record the targets, scoring rules, time gap and exact result, then use Atlas to compare the source passages. Keep missing details open rather than assume the whole process was checked.

Atlas

Check a repeated-rating protocol

Trace the same rater's judgments and keep reported limits with citations.

What intra-rater reliability measures

The repeat check is within one rater. The Wiley reference describes reproducibility under the same experimental conditions. The scope of those conditions matters to the claim.

Imagine one researcher classifying the same saved behaviors in two rounds. The result concerns that person's judgments across rounds. Adding a second researcher creates a question about how two people agree, which needs its own result.

The orthopaedic review separates repeat trials by one examiner from ratings by different examiners. A paper may report both. Keep the values separate: one person may score cases in a stable way even when a team using the same rules does not agree.

If people fill in a form again, the wider question about scores over time is covered by test retest reliability. Find the unit being repeated before deciding which kind of evidence you need.

Identify the procedure being repeated

Suppose a researcher scores a stored recording twice. The behavior in it stays fixed while the researcher judges it again. This checks the scoring part of the process. It does not test how well a new recording would capture the behavior.

Now suppose the researcher records the person again and scores that new clip. The person's behavior, the way it is recorded and the scoring may all vary. The result covers more of the process, so its meaning is different.

The original COSMIN study distinguishes repeating the whole procedure from scoring fixed images again. Table 4 asks which instrument, version, trait and population the study concerns. It also asks which measurement components were repeated and which sources of variation were studied.

COSMIN Table 4 specifying instrument version, repeated measurement components, sources of variation and population. Mokkink et al. 2020, CC BY 4.0

Table 4 is reproduced from Mokkink et al. (2020) under CC BY 4.0. The source PDF was cropped to the table, whose contents are unchanged.

Those fields define the claim. The table's repeated components and sources of variation help distinguish scoring fixed material from collecting it again. Its instrument version, construct and population fields keep the result tied to the right tool, trait and group.

Find each part in the paper's methods text. If the rules or training changed between rounds, retain that detail too. The COSMIN account helps frame what can vary, but the actual study must tell you what it held fixed.

Check independence and the rating interval

Record how much time passed and why. A gap may help reduce recall, yet it does not show that the rater could not see old scores. Time and access to old results are two separate facts.

Look for a statement that earlier scores were hidden. A new order for cases and a new score sheet may also matter. Keep what the study did distinct from the effect it hoped to achieve.

The specialist review discusses concealing previous results where possible. It also explains why the break depends on the task. Avoid assigning one ideal gap to every kind of rating study.

If the paper says only that the rater scored cases again after two weeks, preserve that fact. It does not establish that recall bias was gone. For live examinations, also ask whether people changed during the gap; for stored clips, ask whether the material and display stayed the same.

Read the reported agreement statistic

The score scale helps define which question the result answers. Labels with no order, ordered labels and numerical scores pose different questions. The Wiley reference discusses approaches for these outcome types. Keep the chosen method tied to the scale and design.

For an ICC, retain the model and whether it concerns one score or a mean of scores. Also retain the technical choice between consistency and absolute agreement. Koo and Li explain why these forms answer different questions.

For kappa-family results, copy the exact statistic and any weighting. A raw percentage agreement and a chance-adjusted coefficient describe different things. Do not swap their names just because both concern how ratings agree.

Keep the confidence interval next to the value. Koo and Li use this uncertainty in reading results. A value alone leaves the range of plausible results out of view.

A single pass threshold is rarely enough to guide use. Ask what disagreement would mean for the planned task. Also keep accuracy of the intended trait separate from stable scoring: a rater can repeat the same mistaken judgment.

A worked repeated-rating evidence note

Imagine a made-up plan in which one observer scores 12 stored videos twice with the same rubric. The second round is ten days later and uses a new video order. The plan names an ICC but omits its form and says nothing about access to earlier scores. There are no real ratings or computed results here.

The note keeps the stated procedures beside what still needs checking. It lets a reader see that the saved videos stayed fixed, without assuming that collecting new videos would give the same result.

DetailProtocol evidenceRemaining question
RaterSame observer in both roundsHow relevant is this person to intended use?
MaterialSame 12 stored videosThe recording process was not repeated
Repeat procedureSame rubric, ten-day gap, changed orderWere earlier scores hidden?
ResultICC named without modelWhich form and uncertainty were reported?

Table 1: For a real paper, replace the teaching details with source pages or sections. Keep each rater's own repeat result separate when several people take part.

The assessment chapter gives wider context for reliability types; use the study plan to settle which design applies here.

Trace repeated judgments in Atlas

Add the permitted plan, rubric and results paper to one project. Once you can read the sources, start a chat. Use @ in Ask a question to select them, then ask:

Trace this same-rater repeat procedure. Find the targets, repeated parts, scoring rules, time gap, order, access to old scores and reported statistic. Cite each fact. Mark missing facts as not reported. Keep intra-rater and inter-rater results separate. Do not generate ratings or calculate agreement.

Select Send and open each numbered citation. Check whether the passage says old scores were hidden or only names the gap. If the answer treats these as the same fact, correct it and retain the distinction in your note.

When an answer blends the plan with a later paper, ask what each source contributes. A citation may point to a related section that leaves your question open. Find the named source yourself and read the relevant text if the citation is unhelpful.

Select New, then Note, and save the corrected note with source pages and missing facts. Wait for Saved before closing. Atlas helps with reading and comparison. Researchers still make the ratings, choose methods, compute results and judge what the process can support.

Preserve the limit of the claim

Name the rater, targets, rule version, repeated part, time gap and exact statistic. These facts show what stayed fixed and what varied. Keep missing proof about access to old scores in the claim's limits.

A result for scoring fixed videos again does not cover new recordings, new raters or a changed rubric. New test versions concern parallel forms reliability; links among items concern internal consistency.

Return to the note when a supplement or new scoring guide adds detail. Revise the claim and keep a record of why its earlier scope was narrower.

Atlas

Check a repeated-rating protocol

Trace the same rater's judgments and keep reported limits with citations.

Frequently Asked Questions

It concerns consistency in repeated judgments by the same rater on the same targets under specified conditions. Interpret the result with the repeated procedure, material, interval and statistic.