Skip to main content

Blog

Conventional Content Analysis: From Text to Categories

Learn conventional content analysis with a worked coding memo. Trace each category to text, revise weak labels, and keep researcher decisions visible.

Semantic Map: Visualize the topic from new angles.
Knowledge Map: Deconstruct the article into its structure.

Conventional content analysis builds codes and categories from the text you study. A code is a short label for an idea in the text. A category groups codes that deal with a shared concern. Read the accounts, label the ideas, and check how they relate. The starting groups come from the data rather than a theory you chose in advance.

The hard part is showing how you got there. A useful coding memo links a group to the text behind it, states what fits, and records why you changed your mind. The worked case below shows a researcher narrowing a label when a comment that seems to fit has a different meaning.

Atlas

Check the passages behind your categories

Compare selected text, inspect citations, and save a checked coding memo.

Start with categories from the text

Hsieh and Shannon's 2005 paper distinguishes conventional, directed, and summative content analysis by the origins of their codes and how analysis proceeds. For the conventional approach, categories develop from the material. It suits a descriptive question when existing theory or research offers limited guidance about the phenomenon.

That choice does not mean reading without a question or any prior knowledge. You still decide what you want to understand and which material can address it. Keep a note of beliefs that might lead you to favor one reading. The practical guide by Erlingsson and Brysiewicz treats analysis as reflective work, with repeated returns to the original text.

A topic label such as “student support” is too broad to show what a group means. You might instead group accounts about keeping advice for later use. Find the text that supports this reading and explain why the codes belong together.

Then check the group against the rest of the text. Giving it a name is one step in that work.

Identify where the starting codes come from

Conventional analysis develops the initial labels

The original conventional-analysis account starts with whole-data reading, close attention to relevant words, and notes that develop into codes. Related codes are then grouped and defined.

Keep that origin visible when describing your method: what did you read, which ideas did you label, and how did you decide they belonged together?

For the help-desk case, “keeping advice for later” could be a code drawn from a student's comment. The researcher must check whether other accounts describe that same concern. A heading in the interview guide alone does not show that the accounts belong in one group.

Directed analysis uses a starting framework

In a directed approach, theory or prior research guides the first codes. The OHSU record of the three-approach paper states that distinction explicitly.

If you begin by assigning every comment to a model's fixed categories, report that starting point. Finding a new code later does not erase the role the model played at the start.

You can compare findings with a theory even if it did not supply your first codes. State when you used the theory and what it did. A model might help you discuss a result after coding, or it might set the codes from the start. Those roles need their own account in the method section.

Summative analysis starts with counts and context

Summative analysis counts and compares selected words or content, then interprets their context. The open textbook's approach comparison explains that sequence.

Counts can also appear in a conventional study, but the presence of a frequency table does not determine where its codes came from.

Two people may use the same word for different reasons. In the example below, “email” appears in two comments, yet one concerns a record and the other concerns access to a person. Counting the word alone would hide that difference.

Zhang and Wildemuth's qualitative content chapter emphasizes meanings in context rather than treating occurrence as statistical importance. Explain the job of any counts you report and the unit being counted.

Prepare a corpus you can return to

State which material answers the question

Choose material for a specific question and record its boundaries. For a help-desk project, that could mean permitted interview transcripts about new students' first attempts to get course advice.

Record who supplied the text, the period it covers, and what was left out. Those details let a reader assess the reach of the account.

Elo and colleagues' trustworthiness paper connects data collection, sampling, and unit choices to the study's purpose.

If you selected only people who successfully reached the help desk, do not write as though you captured the experience of every student who needed help. A useful category cannot repair a source set that misses the question.

When the corpus consists of published papers or documents, retain editions, dates, and selection reasons too. The literature review process can help you organize that broader search and selection work.

Coding what documents say and synthesizing research findings are different jobs; explain which one your project is doing.

Keep the reason with the action

A meaning unit is a piece of text that conveys a relevant idea. Its size depends on the passage and the question. In the example, keep “so I can check it before enrollment” with “I ask for an email.” Without the reason, a code may describe only the channel and miss what the student wanted from it.

The trustworthiness paper's unit guidance warns that overly broad units can contain several meanings, while overly narrow units fragment them.

Read nearby text before selecting the unit. If one account shifts from finding a person to keeping their advice, two units may better preserve the distinction than a single channel label.

Preserve the original beside your working notes

Keep source locations with any shortened wording and provisional code. A short note is useful only while you can return to the account it represents. The hands-on guide's condensation example shows shortening while retaining core meaning.

Check that a summary has not added a motive or removed a condition before using it to group passages.

Give each note a source label and a place to return to, such as the transcript name and page used in your project. If you change the wording, record why. Keep these details in your own study records. Atlas source comparison does not replace the coding database or change log your project requires.

Build and revise a coding memo

The three comments below are original teaching examples. They are invented student statements, not interview data. The fictional question is: “How do new students describe getting and using course advice from a help desk?” No participants were recruited and no findings are claimed.

Label what each passage says

Read each invented account with its reason intact. The initial codes stay close to the stated concern. Their different meanings matter even when two comments mention email.

Invented passage and locationProvisional codeWhat supports the code
A, passage 1: “I ask them to email the advice so I can check it again before enrollment.”Keeping advice for laterThe student wants to return to the advice at a later task
B, passage 1: “I wrote down the answer, but later I could not tell which course it referred to.”Losing the link between advice and courseA record exists, but the student cannot apply it to the right course
C, passage 1: “I send an email because the desk is closed when my shift ends.”Reaching staff outside desk hoursThe stated reason concerns when staff can be reached, rather than keeping advice

Table 1: The three invented passages support different codes because their stated reasons differ. The practical guide's coding discussion uses descriptive labels close to the text and allows revision. For this example, “lack of trust” would add a motive that passage A does not state. The student may trust the advice and still need to read it again later. Retain the narrower code unless the surrounding material supplies that claim.

Group the shared concern without hiding differences

An early category called “preferring written advice” seems to link all three comments. A closer read shows the poor fit. B does not say which form of advice the student prefers. C says why the student sends email: the desk is closed. The accounts share a channel, but the question asks how students get and use advice. That needs a check of what each person means.

A revised category could be keeping advice usable later. A describes getting a lasting record; B describes a record that loses its link to a course. Their codes concern two sides of whether advice stays useful. C falls outside this group. Its code may belong in a group about reaching staff, once the wider text supports that choice.

This grouping is a proposed reading of three invented comments. In an actual project, test it against all relevant material and competing readings. The open textbook's conventional process describes grouping related codes and defining categories; the example here shows why the grouping needs a reason beyond shared vocabulary.

Write the definition and the revision reason

A short working memo could read:

Category: Keeping advice usable later. Include text about keeping advice or the details needed to use it at a later task. A concerns reading the advice again; B concerns knowing which course it fits. Leave C outside because the stated reason is reaching staff when the desk is closed. Changed from “preferring written advice” after checking why C uses email. This is a working definition. Check it against the rest of the text.

The memo states what the group means, what fits, which text supports it, and why the name changed. It leaves room for a further choice. A new account might show that losing the course name belongs in a group about understanding advice. Keep that option open and check it against the text before you settle on the first grouping.

The coding-manual discussion by Zhang and Wildemuth supports definitions, examples, and evolving notes. Your project's format may be more detailed, but a label without those decisions gives another reader little basis for checking your interpretation.

Check what the categories leave out

Look for cases that challenge the definition

Return to material that fits poorly, including accounts that use different words for a similar concern. A student might say “I saved the message” or “I took a photo of the form.” The relevant question is whether the passage describes keeping advice usable later, not whether it contains “email.”

Elo and colleagues discuss category coverage and the degree of interpretation in analysis. A contrary passage can prompt a narrower definition, a split, or a different grouping.

Record which option you chose and why. You do not have to make every passage fit the first label.

Separate examples from corpus claims

Three clear excerpts can show what a category means without showing how well it fits the full source set. Read the rest of the text and record cases you are unsure about. Does the label leave out a different experience? Explain how the quoted accounts relate to the wider work.

The original paper identifies context loss as a challenge. Keep enough nearby text in view to avoid that loss. Read the cited words with the rest of the account to check what they support.

One quote does not show that you have covered the full source set, reached saturation, or checked agreement among coders.

Choose checks for the study design

State who reviewed the work and what they checked. If the team disagrees about whether B fits the group, note that question and return to the text together. A clear label can still hide a dispute about what the words mean. Check the meaning before you agree on the name.

The trustworthiness paper discusses researcher review and reporting. An AI answer is not an independent coder.

Choose and report your checks with the methods team. This guide sets no single agreement percentage or sample size for every study. It runs no test of coder agreement.

Compare selected passages in Atlas

Use Atlas for a bounded comparison between selected text and your own memo. Keep your corpus records and coding decisions in the system your project requires. Source-grounded reading can help you inspect a rationale while you retain responsibility for the analysis.

  1. Add permitted source material and your initial coding memo to a project. Confirm that each item has finished processing.

  2. Start a chat. Type @ and select the sources and note to compare. Choose Project only to prevent new web or literature retrieval. Earlier chat context remains available, so name the passages and memo you intend to use.

  3. Ask: “Compare these selected passages with my category definition. Give supporting citations, identify wording that adds an unstated motive, and keep cases that may fall outside the category separate. Do not claim you have checked the whole corpus.”

  4. Open each numbered citation and read the passage with nearby text. Check the source identity and whether the reason for the action is retained.

  5. Correct the comparison. If it labels passage A “distrust,” remove that inference and retain the stated wish to check advice before enrollment.

  6. Choose New, then Note. Save your revised definition, source locations, change reason, and unresolved cases. Wait for Saved before closing.

The capture below shows an actual Atlas source beside a cited answer. The visible ColPali paper is unrelated to these invented student comments. It illustrates where to inspect a passage, without claiming a live content-analysis test or a measure of coding accuracy.

ColPali paper beside an Atlas answer and numbered source citations

Read the surrounding text before keeping the answer's interpretation in your memo.

The displayed paper is Faysse and colleagues' ColPali, dedicated to the public domain under CC0 1.0. It is displayed unchanged within the Atlas capture.

If a citation is missing, ask for support from the named item. If it will not open, find the source by name and inspect it directly. The AI citation checker workflow develops this check between claim and source. An answer that cannot support a category claim needs correction before you save it.

Report the route from text to findings

Keep provisional notes separate from findings

The worked memo gives a starting reason for a group. A final report must describe the actual source set, how you coded it, what changed, and what you checked. Then explain what you found. Do not write that “students prefer email” just because two teaching comments mention it.

The trustworthiness reporting guidance calls for a visible connection between data and results. Give readers the context needed to judge your account, including what the category leaves unanswered. Keep claims within the material and design you actually used.

Compare the findings with prior work

Once your categories are supported, compare them with relevant literature. Keep the contribution of each source distinct. The research synthesis workflow helps with that broader comparison, while the coding memo retains the path from your own selected material to its category.

A reader should be able to follow why “preferring written advice” became “keeping advice usable later,” locate the passages behind the change, and see why the outside-hours comment stayed separate. That is the work the memo makes reviewable before you turn it into a finding.

Atlas

Check the passages behind your categories

Compare selected text, inspect citations, and save a checked coding memo.

Frequently Asked Questions

It is a qualitative approach that derives coding categories from the text data. Researchers read the material, develop labels, group related codes, and check their account against the original text.