Skip to main content

Evaluation Plan: A Question-to-Evidence Template

Build an evaluation plan that links decisions, questions, criteria, data sources, methods, timing, and reporting. See a worked question-to-evidence matrix.

Byline
Jet New
Jet New

Summary

  • An evaluation plan says what a program needs to learn, who will use the findings, and how each question will be answered.

  • Link each question to a success rule, data source, method, time, owner, and report audience before collection starts.

  • The worked matrix flags a food-access program's missing baseline and a survey that cannot prove its reach claim.

An evaluation plan says what a program needs to learn, who will use the answer, and how the team will get sound evidence in time to act. It ties each question to a success rule, source, method, date, owner, and report.

AHRQ's plan guide treats those parts as linked choices.

Here is a small test: if a food-access program wants to know whether referrals reached residents, a survey of workshop attendees cannot show the reach rate for all referrals. The plan needs a referral list and a clear rule for a completed contact.

Atlas can compare the program brief, draft plan, and past reports with cited passages. The team must still choose the right measure and collect the missing data.

Atlas

Check your plan sources in Atlas

Compare the program brief and draft plan with cited passages.

What an evaluation plan includes

An evaluation plan is a written agreement about purpose, questions, criteria, data, methods, schedule, roles, analysis, and use. Its point is to make a future finding both useful and feasible.

The National Institute of Justice says a plan should be made with people involved in the program, ideally before it begins.

It should say who will decide what after reading the findings. A manager may need to change how a service runs next month. A funder may need a year-end result. Those uses can call for different questions and dates. AHRQ starts with the decisions and when they must be made.

The plan is also a check on weak assumptions. A record of workshops delivered can answer whether work took place. It cannot by itself tell you whether residents used a service or whether the program caused a later change. The data and design have to match the claim.

Start with use and decisions

Name the main users of the findings and the action each may take. Ask when that action is due. A plan for midcourse improvement may need quick process feedback. A final outcome review may need a baseline, follow-up data, and more time.

The CDC's current training plan guide links evaluation purpose to the decisions it should support.

Ask staff, partners, and people served what they need to know. They may define success in different ways. If they do, record the difference before picking measures.

AHRQ's guide says shared criteria make findings more credible to stakeholders.

Set scope that the team can fund and run. Name the program sites, dates, groups, and questions included. If one site uses a different service model, note it.

The NIJ planning guide warns that multi-site work may need site-specific records and designs.

Turn objectives into answerable questions

Read the program aim, grant terms, and a simple account of how the work is meant to help. A program theory can make the expected path clear. Turn a broad aim such as "improve access to healthy food" into questions that can guide real choices.

One question might ask, "What share of eligible referrals led to a completed food-service contact within 30 days?" Another might ask, "What stopped residents from completing a referral?"

The first needs a defined group and linked records. The second may need interviews with people who did and did not complete the process. CDC separates program-level evaluation questions from individual survey items.

For each question, write the rule for a useful answer. What counts as an eligible referral? What is a completed contact? Which groups and dates will be compared? A survey item is a tool, while the question and rule tell the team what the tool must help answer.

Ask whether the desired answer is about process, outcome, or cause. Process questions concern delivery and reach. Outcome questions concern change. A claim that the program caused the change needs a design that can address other reasons for it.

Brown University's guide distinguishes process and outcome uses, while NIJ urges early design work for outcome claims.

A worked question-to-evidence matrix

This fictional teaching plan concerns a city program that trains clinic staff to refer residents to food services. Its aim is to make referrals work better.

The table connects each possible decision to a question, success rule, data source, method, date, and owner. It does not report real results.

Diagram linking a fictional food-access decision to an evaluation question, record check, date, and owner, with a stop sign for a missing baseline
Original teaching diagram for the fictional program. The referral question has a record-check path and owner. The wider food-access claim needs a baseline and a suitable comparison before it can be answered.
Decision and questionRule, source, and methodDate, owner, and gap
Fix the referral handoff? What share of eligible referrals led to contact in 30 days?Define eligible referrals and a completed contact. Link clinic referral logs with service intake records; check missing IDs.Monthly review by clinic and service data leads. Gap: the intake file may not use the same ID. Test linkage before promising a rate.
Change staff training? Where do staff and residents say the handoff breaks down?Interview staff and a sample of residents, including people whose referrals did not lead to contact. Compare themes with the logs.First review after eight weeks, led by an independent evaluator. Gap: recruitment and consent need a plan.
Claim better food access? Did residents' access improve because of the program?A workshop feedback survey does not measure access or support a causal claim. Baseline and follow-up measures plus a credible comparison would be needed.Design decision before launch. Gap: no baseline or comparison is in the fictional brief; narrow the claim or collect suitable data.

Table 1: The first row can be useful even if the team later drops the causal question. It tests whether the two record systems can be joined.

If they cannot, the team can report separate counts and disclose the limit. It should not divide one system's completed contacts by the other's referrals as if every record matched.

The last row is a deliberate stop sign. A positive survey from people who attended training cannot tell us how food access changed for residents who were never surveyed. The team can narrow the question to what it can observe, or fund a sound design before making a wider claim.

AHRQ warns that an easy count, such as reports sent, may say little about whether people used the information.

Match methods, timing, and owners

Choose a data source and method for each question, then test whether the team can use them. Name the people and units in the data. Check access, missing records, consent, and who can collect and read the data.

A method that sounds rigorous on paper is of little use if the source will arrive after the decision date.

Plan early for a baseline when change over time matters. A baseline is a measure taken before or near the start of the work. If the program has already begun, say what can still be measured and what was lost.

NIJ notes that asking people later to recall earlier views can bias a comparison.

Match the burden to the purpose. A short process check may use existing records plus a few interviews. A strong claim about effect may need a different sample, more time, and skilled help.

CDC's guide names observation, tests, surveys, review, interviews, and groups as options. Feasibility and timing must shape the choice.

Assign an owner and deadline to each task. Someone must secure data access, test IDs, draft questions, recruit participants, check consent, analyze results, and prepare the report. The KU Community Tool Box stresses that a useful plan also needs a realistic timeline and resources.

Plan analysis and reporting early

Decide how each source will answer its question before collection begins. For referral records, specify the rate, exclusions, missing IDs, and group breakdowns. For interviews, specify how notes will be reviewed and how differing views will be kept. Avoid adding a method to the plan without naming how its data will be used.

Set the date and form of each report around the choice it serves. A service manager might need a brief monthly handoff note. A partner board may need a longer annual review with limits and changes since the last plan.

AHRQ asks who will prepare findings, who will act on them, and which format each audience can use.

Write a rule for revisions. If the referral system changes, record the date, reason, and effect on the measure.

The NIJ planning article advises teams to revise plans when goals or data access shift and to keep the reasons for each change. A version history helps readers see which results remain comparable.

Next step: check plan assumptions in Atlas

Add the program brief, grant terms, prior reports, and draft plan to one Atlas project. Mention the selected sources in chat. Ask: "For each plan question, show the stated program aim, promised measure, available source, and any conflict. Cite the passage behind each point and mark missing support."

Open every cite and read its context. If a grant asks for a resident outcome but a draft row uses only attendance records, mark the mismatch. If an old report names a baseline that the current plan cannot access, verify it with the data owner before adding it to the design.

Save a checked matrix and open questions in a note. This follows the same source-to-claim discipline as research synthesis.

Atlas helps compare supplied text and find passages to check. It cannot grant data access, choose a valid comparison group, obtain consent, or know whether a measure works at a site. The evaluator and stakeholders make those choices and sign off on the plan.

Resolve gaps before approval

Check every row for a question that the source cannot answer, a measure that lacks a clear rule, or a deadline that falls after the decision. Check who is missing from the sample. A survey of attendees alone cannot show the experience of people who never joined the program.

If the baseline is gone, narrow the change claim or use a defensible design with its limits made plain. If records cannot be linked, change the measure before a rate is reported. If consent or data access is unresolved, do not promise a method that depends on it.

Brown and AHRQ both tie design and data collection to the question's real use.

A plan can include later method-specific work. For example, outcome harvesting can help find unexpected actor changes once the team knows who will use those findings and how they will be checked. The plan should say where that work fits, who will conduct it, and what it can establish.

Atlas

Check your plan sources in Atlas

Compare the program brief and draft plan with cited passages.

Frequently Asked Questions

It is a written plan for why a program will be evaluated, who will use the findings, what questions matter, and how and when evidence will be gathered, analyzed, and reported.