How to Measure Whether a STEM Programme Is Actually Working
Attendance proves that learners entered the room. It does not prove that they became more capable. Strong evaluation follows what students can design, explain and transfer.

The central idea
Measure implementation, learner capability, inclusion and sustainability together; no single headline number can represent programme health.
Editorial evidence note
This article provides professional educational guidance. Any illustrative school situation is hypothetical unless a named external source is supplied.
A practical Ghanaian school scenario
A school team facing this decision could begin with one learner group and one term. The team would define the intended capability, document current constraints, test the approach represented by “Track delivery fidelity and time-on-task”, and review learner work with teachers before expanding. The scenario is intentionally hypothetical so that schools can adapt it without mistaking it for a reported InovTech outcome.
The decision beneath the headline
For schools, NGOs, funders and programme managers, this question has consequences far beyond a single lesson or purchase. STEM reports frequently count workshops, kits and photographs because they are easy to collect, even when those measures reveal little about learning quality.
Measure implementation, learner capability, inclusion and sustainability together; no single headline number can represent programme health. That standard helps institutions distinguish visible activity from durable educational value.
Track delivery fidelity and time-on-task
Evidence should shape track delivery fidelity and time-on-task from the beginning. Define a baseline, preserve learner artefacts, observe the quality of reasoning and decide which result would trigger adaptation rather than expansion.
When teams “Write a theory of change”, they should document both the result and the conditions that produced it. That discipline prevents a successful demonstration from being mistaken for a sustainable programme.
Use performance tasks to reveal knowledge transfer
“Use performance tasks to reveal knowledge transfer” should be translated into a visible decision, not left as an aspiration. For schools, NGOs, funders and programme managers, that means naming the learner behaviour, adult responsibility, resource requirement and evidence that would show the decision is working.
A useful stress test is to attempt “Choose a small balanced indicator set” with the smallest realistic group. Record where time, confidence, access or coordination breaks down; those observations are design evidence, not reasons to abandon the ambition.
Measure teacher independence and confidence
The case for measure teacher independence and confidence becomes stronger when teams separate educational necessity from attractive extras. Begin with what learners must understand or perform, then work backward to tools, staffing and timetable.
In practice, “Collect a baseline” creates an early checkpoint. It gives leaders something concrete to examine before scale makes weaknesses expensive or difficult to reverse.
Disaggregate participation and technical roles
Implementation often fails at the handover between a good idea and ordinary school routines. Disaggregate participation and technical roles must therefore appear in lesson preparation, role descriptions, budgets and review meetings—not only in the programme proposal.
Use “Review evidence during delivery” as an ownership test: identify who acts, by when, with which resources, and what happens if the assumption proves wrong. Clear ownership protects both quality and trust.
Combine numbers with artefacts, observations and learner voice
Equity changes the meaning of combine numbers with artefacts, observations and learner voice. Ask who receives meaningful technical time, who is asked to document rather than build, whose language or disability creates friction, and whether the design quietly rewards learners who already have access.
The action “Use findings to improve, not merely report” should be reviewed with learner and teacher voice. Participation figures alone cannot show whether people experienced belonging, intellectual challenge and genuine responsibility.
A disciplined implementation sequence
Begin with the smallest version that can still test the central claim: measure implementation, learner capability, inclusion and sustainability together; no single headline number can represent programme health. Protect time for preparation, observe what participants actually do and review evidence before adding more learners, locations or technology.
The sequence below converts the argument into accountable work. It is intentionally concise so a school or programme team can assign owners and dates during one planning meeting.
- Write a theory of change
- Choose a small balanced indicator set
- Collect a baseline
- Review evidence during delivery
- Use findings to improve, not merely report
Frequently asked questions
What is the most important starting point for impact & evaluation?
Begin with a clearly defined learner or institutional outcome, then assess people, time, infrastructure and evidence before choosing tools.
How can a school apply this guidance?
Start with a contained pilot, use the article’s action checklist, collect evidence from learners and teachers, and improve the model before scaling.
Put the article into practice
- 1Write a theory of change
- 2Choose a small balanced indicator set
- 3Collect a baseline
- 4Review evidence during delivery
- 5Use findings to improve, not merely report