Skip to main content

Blog

Levels of Evidence: Classify Designs With a Named Framework

Use levels of evidence with a named framework and research question. Trace designs in real papers, record the starting level, and keep appraisal limits clear.

Semantic Map: Visualize the topic from new angles.
Knowledge Map: Deconstruct the article into its structure.

Levels of evidence help you locate a study design within a named framework. To use them well, keep the framework, question, methods passage, and limits with the number.

Start with a framework-design-level evidence note. A label such as “Level 2” loses meaning when the scheme or question is missing. The example below traces real reports to Oxford's 2011 rules. It is a design-matching exercise, not a full appraisal.

Atlas

Trace evidence levels to their framework

Compare the framework and study methods, then save a checked evidence note.

What do levels of evidence tell you?

A level ranks a study design within a scheme built for a stated purpose. There is no single set of numbers for every field or question. Name the framework before using its level.

The Oxford 2011 introductory document helps you find likely useful evidence when time is short. The number does not prove that a study was done well or tell you which treatment to use.

In its treatment-benefits row, an individual randomized trial starts at Level 2. For prevalence, a local, current random survey or census starts at Level 1. The questions differ. A single treatment pyramid would miss that point.

Read the original Oxford table with its introduction and footnotes. The starting match can change after you check how well the study was done, how it fits your question, and how clear its results are.

Choose the framework version and question

Record the framework title, version, source, and question. Use the scheme required by your review or assignment. If none is specified, resolve that choice before classifying the papers.

Keep the version with the level

Oxford's website hosts both 2009 and 2011 resources. Their labels and rules differ. Use the 2011 resource and its version 2.1 table together; do not copy sublevels from an older scheme into it.

Burns, Rohrich, and Chung's methods review explains why the same study can receive a different level under another system. Keep a copied author or journal level separate from your own framework match.

If two instructions name different versions, keep the conflict visible. A document comparison workflow can help locate the difference. The review team or supervisor must settle which rule applies.

Choose the row that answers your question

Oxford has rows for prevalence, diagnosis, prognosis, treatment benefits, common harms, rare harms, and screening. Choose the question first. Do you want to know how common a problem is, how a test works, or whether a treatment helps?

The full source table below shows those question rows against Levels 1–5. The prevalence row favors a local, current random survey; the benefits row places an individual randomized trial at Level 2. Footnotes allow changes after appraisal. Some cells are not applicable.

Original Oxford 2011 levels of evidence table with question-specific rows, Levels 1–5, and grading footnotes

Original OCEBM Levels of Evidence Working Group, Oxford 2011 Levels of Evidence, version 2.1. Full unchanged PDF-page capture. Oxford resource license: CC BY 4.0.

The footnote lists study quality, imprecision, indirectness, inconsistency, and effect size among reasons a level can change. Here, indirectness means the study may not fit your question. Check who was studied, what was done, what it was compared with, and which outcome was measured.

Trace four design matches through original reports

This note uses the Oxford 2011 prevalence and treatment-benefits rows. Each number is a starting design match before appraisal. The examples do not compare clinical treatments or estimate new effects.

Keep separate sample streams separate

Menachemi and colleagues' Indiana report includes two sample streams. One was selected at random in April 2020; the other was nonrandom in May. A shared paper title does not give them the same design.

Sutton and colleagues' Oregon report calls its sample a convenience sample. The facilities were approached in random order and were asked to send random subsets of patient specimens. Those steps do not make the sample random for all state residents.

Original report and questionMethods evidenceOxford 2011 starting design matchLimit to retain
Indiana random stream: prevalence in late April 2020Stratified random selection from resident records; Table 1Level 1 prevalence route for a local survey current to that historical questionLow response, sampling-frame exclusions, and point-in-time scope still need appraisal
Indiana nonrandom stream: prevalence in selected communities in May 2020Nonrandom recruitment and separate analysis; Table 2Level 3 local nonrandom-sample routeDo not inherit the random stream's Level 1 match or state-wide scope
Oregon: seroprevalence in May–June 2020Convenience sample of specimens from health-care settingsLevel 3 local nonrandom-sample routeRandom facility or specimen steps do not remove the convenience-sample boundary
SHAREHD: whether its specified intervention helpsRandomized early/late sequences across 12 renal centresLevel 2 treatment-benefits route for an individual randomized trialCheck conduct, time adjustment, missing data, and outcome-specific support

Table 1: These provisional Oxford 2011 matches preserve each report's sampling design, historical question, and appraisal limits; they do not assign final evidence quality.

Keep each match attached to its sources

For Indiana, read the sampling paragraphs and each table heading. The report's Discussion says the random stream had a low response and cannot generalize to other states or times. Its erratum corrects a results sentence; this note adds no prevalence estimates.

For SHAREHD, the unit assigned at random was the centre, not each patient. The stepped wedge design guide traces this trial's rollout, transition observations, and time adjustment. The design match does not answer whether the reported effect is sound or whether a program should be adopted.

These are historical, question-specific matches. An April 2020 survey cannot become evidence of current 2026 prevalence by retaining a Level 1 label. If your question changes, revisit the scope and framework row.

Check the methods and keep appraisal separate

Read enough of the methods to support the design label. Keep what the report says distinct from what you infer. A title, abstract, or database tag may be a useful lead, but it is not the full design evidence.

Inspect how the study was conducted

Find how people or groups entered the study, what was compared, when measures were taken, and what happened to missing data. Note the page or section for each feature that matters to your rule.

Sargeant, Brennan, and O'Connor's conceptual review sets design levels apart from checks of study features and risk of bias in context. Its veterinary examples show how a design name can miss flaws in what was done or measured.

A source-checking workflow can catch a note that repeats an author label without its methods. When wording and methods seem to conflict, preserve the passages and mark the classification as unresolved.

Burns and colleagues also warn that a method not reported is not necessarily a method not performed. Write “not reported in this source” when that is what you know. Do not turn a reporting gap into proof of bad conduct.

Separate design level, bias, certainty, and recommendation

A design match is one part of reading evidence. A risk-of-bias review asks how the way a study was done could distort its result. Keep that tool and those judgments distinct from the design-level note.

The GRADE Working Group's current requirements address how sure we are about a body of evidence for each outcome. They require domains such as risk of bias and imprecision to be checked. Its certainty categories are not new names for Oxford's five levels.

A recommendation requires further judgment. A design rank cannot settle it. Do not translate “Level 1” into “use this treatment,” or call a quick classification note a completed GRADE assessment.

Check what the classification note supports

Before using a level, check that the note includes the framework version, question row, design passage, starting match, caveats, and open questions. A reader should be able to follow the match back to both the paper and the rule.

Keep source limits close to the number. The selection bias guide helps examine who enters a sample or stays in a study. Naming a random design does not settle those questions.

Use the appropriate question and field. A clinical treatment hierarchy does not automatically rank qualitative inquiry, engineering prototypes, or historical research. A study may serve another purpose well even when it falls outside this framework.

This note covers the papers you chose. A research synthesis workflow can help compare them. A few checked matches do not show that you found all the evidence or completed a formal certainty assessment.

If you are using prior evidence to plan a new study, the evidence based research guide connects a source finding to a proposed design choice and the rationale it still needs.

Keep final appraisal and clinical choices with qualified human judgment. If the rule or methods remain unclear, leave the level provisional and record what would resolve it. Do not invent a number to make the note look complete.

Save a checked evidence-level note in Atlas

Atlas can help compare the named framework with papers you may process. Ask for a cited design match and its limits. Check the output yourself before using it in a review or assignment.

  1. Add the selected papers and the exact framework version to one project. Wait for the sources to finish processing.
  2. In chat, type @ and select those sources. Name your research question and the relevant framework row.
  3. Ask for the design, methods passage, starting level, framework rule, and caveats. Open each citation and correct mismatches or missing support.
  4. Select New, then Note. Save the checked evidence note with the framework, design, starting level, page or section, and open questions. Confirm Saved before closing it.

Try this request: “Using @selected sources and this framework version, match each paper's methods to the rule for my stated question. Keep sample streams separate. Cite the design passage and framework row. Mark starting matches as provisional before appraisal. Do not invent missing methods, assign GRADE certainty, or make treatment recommendations.”

Atlas

Trace evidence levels to their framework

Compare the framework and study methods, then save a checked evidence note.

Frequently Asked Questions

No. Frameworks differ in purpose, question types and numbering. State the name, version and applicable rule with the level. A number copied from one system cannot be used as though it came from another.