Formative and Summative, Kirkpatrick's Four Levels, and Raab et al.'s Four Functional Types
Introduction
Topics 25 and 26 established what evaluation means, why it matters, and the principles that govern good evaluation practice. This topic covers the specific ways evaluation gets classified in the training and extension literature — by when it happens, by what it measures, and by what function it serves. All three classification systems below come directly from the Halim & Ali/FAO (1997) chapter already anchoring much of this course.
๐ฏ Learning Outcomes
- Distinguish formative from summative evaluation by timing.
- Explain Kirkpatrick's (1976) four criteria for evaluating a training programme.
- Explain Raab et al.'s (1987) four functional types of evaluation.
- Map all three classification systems onto each other.
- Cross-connect this topic to Practicals 3 and 9.
๐ Why This Matters
Kirkpatrick's four levels are ordered by increasing difficulty and increasing value: reaction data is easiest to collect but tells you least about real impact, while results data is hardest to collect but tells you the most — meaning many organisations stop at reaction ("did trainees like it?") precisely because it's easiest, even though it answers the least important question.
1. Classification by Timing: Formative vs. Summative
Formative Evaluation
Collecting relevant, useful data while the training programme is still being conducted — used to catch drawbacks and unintended outcomes early enough to revise the programme's structure to better fit the situation.
Summative Evaluation
Conducted at the end of a programme, making an overall assessment of its effectiveness in relation to its objectives and goals.
2. Classification by What Is Measured: Kirkpatrick's (1976) Four Criteria
Donald Kirkpatrick's four-level model remains the most widely cited framework for evaluating a training programme, each level measuring a different aspect of the programme:
| Level | What It Measures | Extension Example |
|---|---|---|
| 1. Reaction | How trainees felt about the programme's content, methods, duration, trainers, facilities, and management | Farmer feedback forms after a KVK training session |
| 2. Learning | The skills and knowledge trainees were actually able to absorb during training | A short knowledge test on a new practice, before and after training |
| 3. Behaviour | The extent to which trainees apply what they learned in real field situations | Observing whether farmers actually adopt the demonstrated practice in their own fields |
| 4. Results | The tangible impact of training on individuals, their job environment, or the organisation as a whole | Measurable change in crop yield or income attributable to adoption |
3. Classification by Function: Raab et al.'s (1987) Four Types
Raab, Swanson, Wentling, and Clark's (1987) FAO trainer's guide to evaluation — already introduced in Topic 25 — offers a second classification, organised around the function each type of evaluation serves in the training cycle rather than its timing or what it measures:
| Type | Function | When Conducted |
|---|---|---|
| Evaluation for Planning | Informs planning decisions — choosing or guiding the development of training content, methods, and materials | Before the programme, during design |
| Process Evaluation | Detects or predicts defects in the procedural design of a training activity, allowing mistakes to be caught and corrected before they become serious | Periodically, throughout implementation |
| Terminal Evaluation | Determines the effectiveness of a training programme after it is completed, including the causes of any failure | Immediately at the end of the programme |
| Impact Evaluation | Assesses changes in on-the-job behaviour resulting from training, gathering feedback from trainees and supervisors about real-world outcomes | Some time after the programme has ended |
4. Mapping the Three Systems Onto Each Other
| Formative/Summative | Kirkpatrick Level | Raab et al. Type |
|---|---|---|
| Formative | Reaction, early Learning checks | Evaluation for Planning, Process Evaluation |
| Summative | Learning (final), Behaviour, Results | Terminal Evaluation, Impact Evaluation |
These three systems classify the same underlying activity — evaluating a training programme — along different dimensions: when it happens (formative/summative), what it looks at (Kirkpatrick's four levels), and what purpose it serves in the training cycle (Raab et al.'s four types). A well-designed evaluation plan typically draws on more than one system at once rather than treating them as competing alternatives.
Cross-reference: Practical 9 (Pre-test/Post-test Evaluation) is a direct, hands-on application of Kirkpatrick's Learning level, using a before/after design to measure knowledge gained. Practical 3 (Planning/Organizing/Conducting) touches the process-evaluation function described above, since ongoing monitoring during delivery is part of conducting a programme well.
⚠️ Reaction Is Not the Same as Results
A programme can score extremely well on Kirkpatrick's Reaction level — trainees enjoyed the sessions, liked the trainer, found the venue comfortable — while scoring poorly on Results, if no lasting change in practice or outcome ever materialises. Reaction data alone should never be mistaken for evidence of genuine impact.
๐พ Extension Angle
A KVK farmer training programme evaluated only through end-of-session feedback forms (Reaction/formative) would miss the more important question of whether farmers actually adopted the demonstrated practice (Behaviour) and whether that adoption produced a measurable yield or income change (Results/Impact) — the kind of evaluation gap that a combined use of Kirkpatrick's levels and Raab et al.'s impact evaluation type is designed to close.
๐ฎ๐ณ Indian Institutional Context
ATMA's reporting requirements (Topic 22) typically demand evidence closer to Kirkpatrick's Results level — demonstrated impact on farmer income or productivity — rather than Reaction-level feedback alone, since continued district-level funding depends on showing genuine outcomes, not just participant satisfaction.
๐ Beyond Agriculture
Kirkpatrick's four levels remain the dominant framework in corporate training evaluation worldwide, commonly summarised in industry as "smile sheets" (Reaction) versus genuine business impact (Results) — the same tension flagged in the misconception box above, just under different terminology.
Frequently Asked Questions
- Kirkpatrick, D. (1976). Evaluation of training. In R. L. Craig (Ed.), Training and Development Handbook. New York: McGraw-Hill.
- Raab, R. T., Swanson, B. E., Wentling, T. L., & Clark, C. D. (Eds.). (1987). A Trainer's Guide to Evaluation. Rome: Food and Agriculture Organization of the United Nations.
- Halim, A., & Ali, M. M. (1997). Training and professional development. In B. E. Swanson, R. P. Bentz, & A. J. Sofranko (Eds.), Improving Agricultural Extension: A Reference Manual (Ch. 15). Rome: Food and Agriculture Organization of the United Nations.