What would change look like? Turning outcomes into meaningful indicators

Wednesday 29th July 2026 by Ciaran McDonald
Illustration representing data points and measurement, reflecting the blog's focus on turning outcomes into meaningful indicators.

This blog is the third in a series of short reflections on quantitative practice by Ciaran McDonald. The previous piece explored why a survey can contain sensible questions yet still generate data that does not answer the evaluation question. This time, the focus is on the step between naming an outcome and measuring it: deciding what the outcome means, what change would look like and what evidence would allow us to recognise it.

Evaluations often focus on outcomes such as confidence, wellbeing, resilience, independence or community cohesion.

These outcomes can sound clear because the words are familiar. But familiarity is not the same as precision. Ask several people what ‘greater confidence’ would look like and you may receive several different answers.

Would someone speak up more often? Feel more able to make decisions? Try unfamiliar activities? Apply for a job? Ask for help when they need it?

All of these could reflect confidence. None captures the concept in full.

Before we can assess whether an outcome has changed, we therefore need to decide what we mean by it in this context – and what we would expect to observe if change occurred.

From a broad idea to observable evidence

In evaluation, abstract concepts such as confidence, wellbeing and resilience are often described as constructs. A construct cannot usually be observed directly in the way we might record someone’s age or count how many sessions they attended. Instead, we infer it from answers, behaviours, observations or records.

Turning a construct into something we can investigate involves two linked processes: conceptualisation and operationalisation.

Conceptualisation asks: What exactly do we mean by this outcome here?

Operationalisation asks: What would we observe if it changed, and how could we capture that evidence?

Imagine an employment programme that aims to increase participants’ confidence in moving towards work. On its own, ‘confidence’ does not tell us enough about the intended change. Here, it might mean greater confidence in explaining skills to employers, completing applications independently or taking part in interviews.

One possible indicator might be:

Participants’ self-reported confidence in explaining their skills and experience to an employer.

A measure could then be a survey item asking participants to rate that confidence at relevant points in their support.

A useful distinction is:

  • An outcome is the change we want to understand.
  • An indicator is an observable sign that provides evidence about whether that change has occurred.
  • A measure is the specific way we collect evidence against that indicator, such as a question, scale, observation protocol or data field.

For an outcome-focused evaluation, the route might therefore look like this:

Evaluation question outcome definition indicator measure evidence claim

Conceptualisation and operationalisation sit at the heart of this chain. Conceptualisation defines what the outcome means in context; operationalisation translates that definition into indicators and measures. Setting out the full chain makes the reasoning visible: how the evaluation question was translated into something measurable, why particular indicators and measures were chosen to represent the outcome, and what the resulting evidence can – and cannot – support.

Indicators shape the story the evaluation can tell

Indicators provide observable evidence about outcomes by representing particular aspects of them. For broad or multi-dimensional outcomes, a single indicator is unlikely to capture the whole picture.

If confidence is measured only through willingness to speak in groups, the evaluation may tell us something useful about confidence in that setting. It cannot automatically support a broader conclusion that participants have become more confident across their lives.

Similarly, if a broad outcome such as wellbeing is represented by one question about feeling positive, the evaluation may overlook other relevant dimensions such as anxiety, day-to-day functioning, social connection or sense of purpose.

The indicators we choose therefore shape which aspects of an outcome become visible in the evidence.

Validity is about whether the evidence around a measure gives us a sound basis for the interpretation we want to make from its results. Two particularly useful ideas here are content validity and construct validity.

Content validity asks whether the content of a measure adequately represents the outcome for the population and context in which it will be used: is it relevant, does it cover the important aspects, and – where people are responding directly – is it understood as intended?

Construct validity asks whether the results behave in ways we would expect if the measure captured the underlying concept.

In simpler terms, the broader question is:

Are we measuring what we think we are measuring?

Useful questions include:

  • Does the measure cover the aspects of the outcome that matter for this population and context?
  • Are any important dimensions missing?
  • Where respondents are involved, can they understand and answer it as intended?
  • Do the results show the relationships or differences we would expect?
  • Are there plausible reasons why the results might primarily reflect something else?

Even a measure with strong validity evidence from previous uses may still be a poor fit for a particular evaluation. We also need to ask whether it can detect change over time in the construct, whether the kind of change expected could realistically become visible within the evaluation period, and whether the measure can be administered, scored or recorded consistently in practice.

These considerations should inform measure selection before data collection begins, as well as how findings are interpreted afterwards.

Three common ways measurement goes wrong

1. A broad outcome is represented by one narrow indicator

A single indicator can work well when an outcome is specific and clearly defined. It becomes more problematic when expected to represent a broad or multi-dimensional concept.

One question asking whether participants ‘feel more resilient’, for example, places considerable weight on a term people may understand differently. Resilience might involve recovering after setbacks, managing stress, adapting plans or seeking support when needed.

A broad question may capture an overall perception, but not tell us which aspects changed or whether respondents interpreted the concept consistently.

The answer is not automatically to add more questions. It is first to decide which dimensions matter for the intervention and evaluation question. Several indicators may be needed for a complex outcome; a narrower outcome may be captured through one focused measure.

Measurement should follow from the outcome, rather than the outcome being squeezed into the space available in a questionnaire.

2. Available data is mistaken for suitable outcome evidence

Evaluations often have access to monitoring or administrative data that is easy to count: sessions attended, activities completed, referrals made, appointments kept or qualifications achieved.

These measures can tell us important things about reach, engagement, delivery and progression. But their availability does not automatically make them evidence of the intended outcome.

Attendance may create the opportunity for change, but on its own it does not show that a separate intended outcome – such as improved wellbeing or employment – occurred. Completing a training course may be a meaningful achievement or intermediate outcome, but does not necessarily demonstrate that someone entered or sustained employment.

Sometimes an available measure is used as a proxy – an indirect stand-in for an outcome that is harder to observe. That can be reasonable, but the assumed link needs to be explicit and credible.

Behavioural or administrative data may be more closely aligned with some outcomes than self-report. Sustained employment, changes in crisis-service use or subsequent recorded offending may all be relevant where the data is suitable and interpreted carefully.

The key question is:

What does this data genuinely provide evidence of?

Good evaluation makes the best use of existing data, but does not assume that data collected for one purpose can support a different conclusion without checking its relevance, quality and limitations.

3. More measures are assumed to mean stronger measurement

When an outcome is complex, the natural response can be to measure more things.

Sometimes that is appropriate. Several well-chosen items can provide broader coverage. Where items are intended to form a scale and function coherently together, combining them appropriately can reduce reliance on any single item and improve the reliability of the resulting score.

But volume is not the same as quality. Five vague or repetitive questions do not necessarily measure an outcome better than one carefully chosen item. Longer measures can also increase respondent burden and reduce response quality.

The aim is to cover the dimensions that matter, at a level proportionate to the evaluation. The right approach might involve one focused item, several related items, an established scale, administrative or behavioural indicators, or a combination of these – with qualitative evidence used alongside them where it can help understand the nature and meaning of change.

The strongest measurement approach is not the one that produces the most data. It is the one that generates enough credible evidence to answer the evaluation question.

What about ‘validated’ measures?

Measures described as ‘validated’ can be extremely valuable. They may come with evidence about whether they capture the intended construct, produce scores consistently and can detect change over time.

However, ‘validated’ is not a universal stamp of quality or suitability. The evidence behind a measure relates to particular versions, populations, settings and uses.

Evaluators still need to ask whether the measure fits the intended outcome, population, context and timeframe; whether it is accessible; whether it can be administered and scored correctly; and whether the burden is proportionate.

Adapting a measure may improve relevance or accessibility, but may also mean that evidence about the original version no longer applies in the same way. The task is to select the measure most appropriate for the intended use – not simply the one carrying the label ‘validated’.

Measurement choices shape the claims we can make

Imagine participants report higher confidence at the end of an employment programme than at the beginning.

There are several possible claims:

  • Participants reported greater confidence in explaining their skills to employers.
  • Participants became more confident.
  • The programme increased participants’ confidence.

These statements are not equivalent.

If the measure asked about explaining skills to employers, the first statement reflects both the scope and source of the evidence.

The second broadens the finding to confidence more generally and drops the fact that the evidence is self-reported.

The third goes further by attributing the change to the programme – a causal claim that cannot be established from the outcome measure alone.

This does not mean findings need to be described timidly. It means the conclusion should match both what was measured and what the wider evaluation design can support.

Precision is not the enemy of a strong impact story. It is what makes that story credible.

What stronger practice looks like

Before choosing a survey item, scale, observation protocol or data field, it is worth asking:

  • What is the outcome, for whom and in what context?
  • Which dimensions of it matter here?
  • What would we expect to observe if it changed?
  • Which indicators, measures and data sources would provide credible and proportionate evidence?
  • What claims would those measures and the wider evaluation design support?

This process makes assumptions visible before data collection begins. It can reveal where an indicator is too narrow, an existing measure is only a proxy, an established measure does not fit, or quantitative measures need to be used alongside qualitative evidence.

A simple test

A useful question when choosing any indicator is:

If this outcome had genuinely changed, what would we expect to observe – and would this indicator allow us to observe it?

Then ask the reverse:

Could this indicator change without the underlying outcome changing?

Referrals might fall because eligibility or recording practices changed rather than need reduced. A self-reported score might also change because participants’ understanding of what ‘good’ looks like has shifted over time, making it harder to interpret the score as straightforward evidence of change in the underlying outcome.

These possibilities do not necessarily make an indicator unusable. They clarify what it means, what else might explain it and what additional evidence may be needed.

Measuring what matters

Outcomes such as confidence, wellbeing and resilience are not impossible to measure. But they are not self-defining.

Good outcome measurement depends on conceptualising the outcome clearly and operationalising it through indicators and measures that represent it credibly.

When those judgements are strong, relatively simple measures can generate meaningful evidence. When they are weak, even sophisticated tools can create an appearance of precision without real clarity.

The indicators we choose shape which aspects of change become visible in the evidence – and, ultimately, what claims the evaluation can credibly make.

In the next blog, Ciaran will explore the question that follows once we have credible evidence that something changed: how do we know whether the intervention made the difference? That means thinking carefully about comparison and the counterfactual: what would have happened without the intervention.