Skip to main content

Evaluation Questions: How to Choose and Test a Useful Set

Write evaluation questions tied to a program and its users. Use a worked question-rationale matrix to check purpose, evidence gaps, and who approves the set.

Byline
Jet New
Jet New

Summary

  • Evaluation questions state what an evaluation needs to learn about a program so its users can make a decision.

  • Choose a small set by checking the program purpose, intended use, available evidence, and data limits.

  • The evaluator and intended users approve the final questions. A draft source comparison cannot validate the design.

Evaluation questions state what a program evaluation needs to learn. They connect the program's purpose with a choice its users must make. A good set is small enough to guide the work and clear enough to check against available evidence.

A fictional library wants to know whether its reading club caused a rise in children's reading scores. It has attendance logs and parent feedback, but no score records or comparison group.

That impact question outruns the data. The team can still ask who attended, what blocked access, and how members used the club.

Atlas can compare selected plans and reports and show the source text behind a draft question. Evaluators and intended users decide which questions matter and whether they can be answered.

Atlas

Check a draft question against program sources

Compare selected plans, check citations, and save a reviewed rationale.

What evaluation questions do

A key evaluation question names an issue the evaluation aims to answer about a program. It sets a boundary for what the team will study. The CDC evaluation questions checklist ties the question to program purpose, user needs, and the time and data available.

A survey item asks one person for one piece of data, such as “How many club sessions did you attend?” The key question is broader: “Who used the club, and what kept other families away?” The survey item may help answer part of it. Switchboard's guide makes this same level distinction.

The question should lead to a choice. A program lead may need to decide whether to change access hours, keep a delivery model, or study an outcome more closely. CDC's program framework puts question focus after context and program description and before evidence gathering. That order helps the team ask what it needs to learn before choosing a data tool.

A set of questions guides the evaluation. It does not itself supply a study design or evidence. The people who will use the findings should review the set before collection begins.

Start with the decision and users

Name the choice

Ask what a named user must decide after the work. “Should the library keep the Saturday club?” is a choice. “Is the program good?” leaves too much open.

Name the people who will use the answer and when they need it.

CDC's checklist asks whether a question is relevant to the program and its users. A funder, a site lead, and a family may need different facts. The evaluator should hear those needs and record any conflict in scope.

Match the program stage

Check what the program is trying to do and how long it has run. A new club may need to learn whether families can join. An established club may also ask about longer-term results.

CDC's framework begins with context and a program description so questions match the work.

Keep the program plan, eligibility rules, and delivery notes near the draft questions. If a goal appears only in an old proposal, ask whether the team still owns it.

A source review workflow can help keep each goal linked to its document.

Ask who is missing

Averages can hide people the program did not reach. Ask whose experience the evaluation needs to include and who may be absent from existing records.

The CDC framework calls for joint work and fair practice throughout its steps.

The users may decide to ask a separate access question. They must also decide how to hear from people safely. A draft list from plans alone cannot stand in for that talk.

Match questions to evaluation type

The method starts by checking the purpose of the evaluation. The same program can be studied for need, delivery, results, or longer-term change. Choose the type that fits the choice and evidence.

Need and fit

A needs question asks whom the program is meant to serve and what gap it addresses. For the library, ask: “Which families want a reading club but cannot use the current times?” That may call for local views and access data. The University of Wisconsin Extension guide maps needs questions to a program's starting situation.

Delivery and process

A process question asks what was delivered and how people experienced it. “How often did the Saturday club run as planned?” can be checked against schedules and logs. “What stopped some families from coming?” needs accounts from those families as well. Eval Academy's type-based examples show why delivery questions are useful during a formative review.

Outcomes and impact

An outcome question asks what changed for members. “What reading practice did families report after joining?” needs a defined measure and a time point. A stronger claim, “Did the club cause higher reading scores?” needs a design that can address other causes. Attendance plus praise cannot answer it.

CDC's updated framework paper links question choice to design and credible evidence. Wisconsin Extension separates outcomes from longer-term impact. An evaluator may use more than one type, but should name the inference each question asks for.

Check relevance and feasibility

Screen each question

Use four checks before the team locks the set:

  1. Purpose: What choice will this answer inform?
  2. Program fit: Which activity, group, or result does it cover?
  3. Evidence: What source or new data could answer it, and who may use those data?
  4. Scope: Can the team analyze the answer within its time and skills?

The CDC checklist calls these concerns pertinence, reasonableness, specificity, and answerability.

A question may sound useful yet fail one screen. Record that gap before promising an answer.

Check the set as a whole

Two questions can repeat the same task. Another can be left out even though a user needs it. Review the set together. The CDC checklist asks teams to explain why topics are included or left out.

Keep the set short enough to fund and report. A question about every activity can scatter the work. A short list that ignores access or harm may also fail the purpose. Ask intended users to rank the choices they need most.

Test the data plan

For each question, name an indicator or type of account, the source, when it can be gathered, and a likely limit.

If attendance logs have no age group, they cannot show which children were reached. Optional family feedback may miss those who left early.

Access matters too. The team may lack consent to use school scores. Do not treat a file as available until the evaluator checks permissions, quality, and fit. CDC's framework places credible evidence after question and design focus for this reason.

Read a fictional question-rationale matrix

This example is fictional. The program, records, names, and reviewer choices were invented to show a planning artifact. They do not report real library results or Atlas output.

A library runs a Saturday reading club for children aged seven to ten. Its plan aims to widen access to guided reading and build a reading habit. The director must decide whether to keep the hours next term.

The team has session logs, sign-up records, and parent comments. It has no school test records.

Draft questionChoice and source basisGap or assumptionHuman review
Who enrolled and attended?Reach; sign-up sheet S1 and session log L1.S1 may omit children who tried to join.Keep, then check access records.
Which families could not use Saturday hours?Schedule choice; parent note P2 mentions work shifts.One note does not describe all nonparticipants.Keep; seek safe outreach.
Did sessions run as planned?Delivery choice; plan A1 and log L1.Log may record attendance but not reading time.Revise log check.
What reading habits did families report?Goal in plan A1; comments P1 to P3.Feedback is optional and self-reported.Keep as a bounded outcome question.
Did the club cause higher test scores?The director asks about impact; no score file or comparison plan.Needed data and design are absent.Defer; do not promise an answer.
Should Saturday hours change?Director's next-term choice; S1, L1, P2.Need views from families who did not enroll.Use as a decision question after access review.

Table 1: The matrix keeps the director's choice beside each question. It also shows where a source is too thin.

P2 can raise an access concern; it cannot tell the team how many families could not attend.

A useful final set might ask who used the club, what blocked access, how the sessions ran, and what families reported about reading habits. The causal score question stays on a future design list. That is a decision to narrow the current evaluation, not proof that scores did not change.

This artifact follows the CDC checklist's relevance and answerability tests. The team must still talk with families and intended users, approve the set, and check data access.

Revise weak questions

A weak question often hides the program, group, time, or intended use. “Was the club successful?” offers no clear test.

Ask instead: “During this term, who used the Saturday club, and what barriers did eligible families report?” That version names the setting and the evidence to seek.

Avoid a causal verb when the design cannot support it. “Did the club improve reading scores?” needs score data and a way to weigh other causes. If those are absent, ask a narrower descriptive question about reported habits. State the narrower scope in the report so no reader mistakes it for an impact finding.

Eval Academy's examples vary questions by the kind of evaluation being done. Switchboard also treats a key question as a guide for later data collection. A survey prompt can come later, after the team knows what the evaluation must answer.

Choose the right question level

A key question sets the main inquiry. A subquestion narrows one part of it. A data collection item asks a person or record for a single fact.

Keeping the levels clear stops a list of survey items from standing in for a useful evaluation question.

LevelLibrary exampleUse
Key questionWho can use the Saturday club, and who cannot?Guides the access inquiry.
SubquestionWhich hours conflict with family work shifts?Narrows one barrier.
Data itemWhich Saturday sessions did you attend?Supplies one part of the evidence.

Table 2: The CDC checklist draws the line between an evaluation question and a survey item. A team may need several items and sources for one key question. Wisconsin Extension's logic-model guide helps place the inquiry near need, process, or outcome.

A lessons learned register can later keep what the evaluation found and where it applies. Its rows should not be used to invent an evaluation question after the result is known.

Choose the final set

Choose a question when an intended user can name the choice it informs, the program link is clear, and the team can gather credible evidence within its limits. Revise a question when the topic matters but its scope or wording is vague. Defer it when the data or design cannot support the answer.

Ask intended users to review this short list. Record who agreed, who had a concern, and which questions remain open. CDC's framework calls for collaboration throughout evaluation. The evaluator owns the final design and must respect the program's ethics and access rules.

Before collecting data, map each approved question to a source, method, time, and way to analyze it. If the map has a blank cell, the team needs a plan or a narrower question. A later change should carry a reason so the report can explain what was and was not answered.

Check question sources in Atlas

Atlas can help compare the program plan and context files the evaluator has chosen. Add only sources the project may hold. Ask for one row per draft question with the linked objective, cited passage, source date, assumption, and evidence gap. Keep the questions marked as drafts.

Open each citation. Check that the quoted goal is still current, the group named in the question matches the plan, and a claimed data source exists. If Atlas misses a source or blends a goal with a result, correct the matrix by hand. Save a note with the reviewed rationale and unresolved gaps. Public Atlas guides support selected-source comparison, citation checks, and saved notes.

Atlas answer beside an open cited paper; the paper is unrelated to the fictional library program

This first-party image shows an opened citation beside an Atlas answer. The visible paper is about AI science. It is not the fictional club plan and does not show Atlas validating an evaluation design.

The evaluator decides whether to consult families, what data may be used, and how to answer each question. Atlas cannot gather survey responses or make that agreement. For help reading source claims, see qualitative data analysis with AI.

Limits of a draft question set

A program plan may leave out informal work or a changed goal. Compare it with current staff and user accounts before treating it as the whole program.

A source-linked draft shows what the supplied text says. The selected sources may still leave needs out.

Some attractive questions require data the team cannot get. Others require a design that would take more time or money. Mark them for later work or narrow the claim. The CDC checklist asks for questions that a team can answer with its actual access and resources.

Be clear about causal language. A change after a program begins may have other causes. An outcome question can describe the change; an impact question needs a design that can test what caused it. CDC's framework paper places question and design choices together.

Finally, the people who will use the findings must be able to understand the answer. Keep the question set focused, document why each question was chosen, and revisit it if the program or user choice changes.

Atlas

Check a draft question against program sources

Compare selected plans, check citations, and save a reviewed rationale.

Frequently Asked Questions

They are high-level questions an evaluation aims to answer about a program for its intended users.