EXT 505 – Capacity Development | Practical 9
M.Sc. Extension Education · Punjab Agricultural University, Ludhiana
Introduction
Practical 3 introduced Kirkpatrick's Level 2 (Learning) and named the pre-test/post-test as its typical tool, without building one. Practical 5 taught you to write measurable objectives specifically so they could be tested. This practical closes that loop: converting objectives into actual test items, administering them before and after a programme, and calculating whether real learning gain occurred.
Learning Outcomes
- Explain the purpose and logic of a pre-test/post-test design.
- Write test items directly matched to an objective's Bloom level and domain.
- Calculate learning gain using simple and normalised gain formulas.
- Identify threats to validity in a one-group pre-test/post-test design.
- Administer tests appropriately for mixed-literacy audiences.
ЁЯФм Why This Matters
Without a pre-test, a high post-test score is uninterpretable — you cannot tell whether participants learned something new or already knew it before the programme began. The pre-test is what turns a single snapshot into an actual measure of change.
The Logic of Pre-Test/Post-Test Design
This is technically a one-group pre-test/post-test pre-experimental design (Campbell & Stanley, 1963) — the same test given before and after an intervention, to the same group, with no separate control group.
Threats to Validity — Know the Limits
| Threat | What It Means | Extension Example |
|---|---|---|
| History | Something else happened between tests that could explain the change | A govt. advisory on the same topic aired on TV during the training week |
| Maturation | Natural change over time, unrelated to training | Farmers gained general seasonal experience regardless of training |
| Testing Effect | Taking the pre-test itself primes participants for the post-test | Farmers remember pre-test questions and specifically look out for those answers during the session |
| Instrumentation | The test itself changes between administrations | Post-test questions are worded differently or easier than the pre-test |
⚠️ Common Error
Presenting a pre-test/post-test gain as definitive "proof" that the training caused the improvement. Without a control group, this design shows association, not rigorous causal proof — a scientifically honest report states this design limitation rather than overclaiming.
Writing Test Items from Objectives
Every test item should trace directly back to a specific objective's Bloom level (Practical 5) — this is what gives the test content validity.
| Objective (Practical 5 style) | Bloom Level | Matching Item Type |
|---|---|---|
| List 4 common pests of tomato | Remember | Fill-in-the-blank or short recall list |
| Explain why crop rotation reduces pest build-up | Understand | Short-answer or MCQ testing reasoning, not just fact |
| Calculate the correct biofertiliser dose per acre | Apply | Numerical problem/scenario item |
| Compare cost-effectiveness of two IPM strategies | Analyse | Case-based comparison question |
| Demonstrate correct sprayer calibration (psychomotor) | Apply (skill) | Observed performance checklist, not a written test |
⚠️ Common Design Flaw
Writing every test item at the Remember level (simple recall) even when the objective was written at Apply or Analyse. This under-tests what was actually taught — if the objective says "demonstrate," a written multiple-choice question cannot validly measure it.
Item-Writing Guidelines
- One concept per item — avoid double-barrelled questions
- Avoid negatively worded stems ("Which is NOT...") unless unavoidable, and highlight the negative clearly if used
- Keep language at the participants' literacy and vocabulary level, not textbook language
- For MCQs, ensure distractors are plausible, not obviously wrong
- Keep the pre-test and post-test parallel in difficulty and format — same number of items, same structure, different surface wording where possible to reduce testing effect
Calculating Learning Gain
Simple Gain Score
Easy to compute and communicate, but doesn't account for how much room a participant had to improve — someone starting near a perfect score has little room to show gain even with real learning.
Normalised Gain (Hake's g)
Developed by Hake (1998) for education research, this expresses gain as a fraction of the room that was actually available for improvement — useful when comparing groups with different starting knowledge levels.
| ⟨g⟩ Range | Interpretation (Hake, 1998) |
|---|---|
| ⟨g⟩ ≥ 0.7 | High gain |
| 0.3 ≤ ⟨g⟩ < 0.7 | Medium gain |
| ⟨g⟩ < 0.3 | Low gain |
ЁЯзо Tool Note
To test whether the average gain across a group is statistically significant (not just a numerical difference that could be chance), a paired-samples comparison of pre- and post-test scores for the same participants is the appropriate test — available as the Paired T-Test Calculator on this site.
Administering Pre- and Post-Tests
Keep Conditions Constant
Same time allowed, same setting, same instructions for both administrations, to isolate the training's effect.
Immediate Pre-Test, Immediate Post-Test
Pre-test just before the session begins; post-test right after it ends, before recall fades or unrelated events intervene.
Mixed-Literacy Groups
Use oral administration, picture-based items, or a facilitator reading items aloud individually — consistent with Practical 1's guidance on illiterate participants.
Anonymity vs Matching
Use a matched but anonymous code (not name) so pre- and post-scores can be paired per person without compromising confidentiality (Practical 1 ethics principle).
Indian Institutional Context
ЁЯЗоЁЯЗ│ Standard Practice
KVKs routinely administer simple pre-test/post-test forms as part of their training reporting to ICAR/ATARI, typically as a short 10-item quiz. ATMA-funded trainings under SREP similarly require a documented knowledge-gain measure as part of programme reporting and impact assessment.
Beyond Agriculture — Same Evaluation Logic
ЁЯФД Transferable Framework
A hospital's infection-control training and a corporate compliance training both use the identical pre-test/post-test logic, gain-score calculation, and validity cautions taught here. This is a general training-evaluation competency, not agriculture-specific.
Mini Case: Testing the Natural Farming Programme (Practicals 4–7)
Continuing the programme designed across the last several practicals — its Day 1 knowledge objectives are now tested directly:
| Objective (Practical 5) | Test Item |
|---|---|
| Explain the four pillars of natural farming | Short-answer: "Name and briefly explain two of the four pillars of natural farming discussed today." |
| Demonstrate correct Jeevamrit preparation | Observed checklist during Day 2 practice, not a written item |
| Metric | Result (40 farmers) |
|---|---|
| Mean pre-test score (out of 10) | 3.2 |
| Mean post-test score (out of 10) | 7.6 |
| Simple gain | 4.4 |
| Normalised gain ⟨g⟩ | (7.6−3.2)/(10−3.2) = 0.65 — medium-to-high gain |
Note for discussion: this design cannot rule out that some gain reflects the testing effect (farmers recalling pre-test topics) rather than teaching alone — worth flagging honestly in the programme report.
❓ Frequently Asked Questions
Click any question to reveal the answer.
No fixed number, but 8–15 items is a common practical range for a single training session — enough to cover the main objectives without taking excessive time away from the actual training, especially for farmer audiences with limited patience for lengthy paper tests.
Using identical items makes comparison simplest but increases the testing-effect risk (Q. above). A common compromise is keeping the same objectives and difficulty but varying surface wording or example numbers between the two versions.
This happens occasionally — possible reasons include guessing correctly on the pre-test, fatigue on the post-test, or genuine confusion introduced by new information. A few individual decreases don't invalidate the overall group result, but a widespread pattern should prompt a review of the test items or the teaching itself.
It demonstrates Kirkpatrick Level 2 (Learning) only. Claiming full "success" also requires Level 3 evidence (behaviour change in the field) and ideally Level 4 (results/impact) — a knowledge gain alone doesn't confirm farmers actually adopted the practice.
Use an observed performance checklist — a rater watches the participant perform the skill (e.g., sprayer calibration, Jeevamrit preparation) against a list of correct-technique criteria, ideally both before and after the relevant training, exactly as with a written pre/post test but observational instead of paper-based.
References (click to expand)
Research Design and Measurement
- Campbell, D. T., & Stanley, J. C. (1963). Experimental and Quasi-Experimental Designs for Research. Houghton Mifflin.
- Hake, R. R. (1998). Interactive-engagement versus traditional methods: A six-thousand-student survey of mechanics test data for introductory physics courses. American Journal of Physics, 66(1), 64–74.
Training Evaluation
- Kirkpatrick, D. L., & Kirkpatrick, J. D. (2006). Evaluating Training Programs: The Four Levels (3rd ed.). Berrett-Koehler Publishers.
- Mager, R. F. (1997). Measuring Instructional Results (3rd ed.). Center for Effective Performance.
Objectives and Taxonomy (Foundational to Item Writing)
- Anderson, L. W., & Krathwohl, D. R. (Eds.). (2001). A Taxonomy for Learning, Teaching, and Assessing: A Revision of Bloom's Taxonomy of Educational Objectives. Longman.
Extension Education Texts
- Ray, G. L. (2005). Extension Communication and Management. Kalyani Publishers.