๐Ÿ“š Academic Toolkit Dr. Davinder Singh

Sunday, August 30, 2026

Statistical Tools for Evaluation

Topic 30: Statistical Tools for Evaluation | EXT 505
EXT 505 · Capacity Development · Block VI · Theory Topic 30

Turning pre-test/post-test numbers into defensible evidence

Introduction

Topics 27–29 covered the types, process, and use of evaluation findings. This topic supplies the quantitative backbone that turns raw data — pre-test scores, post-test scores, adoption rates — into evidence that can actually support Topic 29's instrumental and conceptual uses. These are not exotic tools: the statistical methods below are the ones most commonly used in published Indian extension education evaluation research, and are the same ones you already applied by hand in Practical 9.

๐ŸŽฏ Learning Outcomes

  • Use descriptive statistics (mean, frequency, percentage) to summarise evaluation data.
  • Explain when and why a paired (dependent) t-test is used for pre-test/post-test evaluation designs.
  • Explain when an independent t-test or chi-square test is more appropriate than a paired t-test.
  • Describe how correlation is used alongside these tests in published extension evaluation research.
  • Locate and use this course's own Paired T-Test Calculator tool.

๐Ÿ” Why This Matters

A published Indian Journal of Extension Education study on farmer knowledge gain from a sugarcane cultivation video used exactly three tools — mean, paired t-test, and correlation — to analyse pre-test and post-test data, showing that a small, standard toolkit is sufficient for credible extension evaluation research, not an elaborate statistical apparatus.

1. Descriptive Statistics

Before any inferential test, evaluation data is typically summarised using basic descriptive statistics: the mean (average score across respondents), frequency (how many respondents fall into each category), and percentage (proportion showing a given result, such as adoption of a practice). These numbers alone are often reported directly in a Kirkpatrick Learning-level or Behaviour-level analysis (Topic 27).

2. The Paired (Dependent) t-test

The paired t-test compares two sets of scores from the same respondents — most commonly, a pre-test score and a post-test score from the same trainees — to determine whether the observed difference is statistically meaningful rather than due to chance. This is the single most common statistical tool in published extension training evaluation studies precisely because most extension evaluations use a before/after design on the same group of trainees.

๐Ÿ”— Cross-Reference — Practical 9 and This Course's Own Tool

Practical 9 (Pre-test/Post-test Evaluation) is where you already applied this exact logic hands-on. This course's own Paired T-Test Calculator tool automates the calculation, letting you focus on interpreting the result rather than the arithmetic — the same role a spreadsheet or statistical software plays in the published studies referenced throughout this topic.

3. Correlation

Correlation measures the strength and direction of a relationship between two variables — for example, whether greater exposure to a training video is associated with greater knowledge gain. It does not, on its own, establish that one variable causes the other, but it is routinely used alongside the paired t-test in published extension research to explore what factors relate to observed knowledge gain.

4. When Other Tools Are More Appropriate

ToolUsed WhenExample
Paired t-testComparing the same group's scores before and after trainingFarmer knowledge scores before and after a sugarcane cultivation video
Independent t-testComparing two different, separate groupsComparing knowledge gain between an app-based learning group and a printed-material group of health workers
Difference-of-differencesComparing the size of the gain (not just the endpoint) between two groups(Post-test − Pre-test) for Group 1 minus the same for Group 2, isolating which intervention produced a larger gain
Chi-square testComparing categorical data (yes/no adoption, category counts) rather than continuous scoresWhether adoption of a practice (adopted/not adopted) differs meaningfully between two farmer groups

⚠️ A Real Worked Comparison

A recent study on Anganwadi Workers' refresher training compared a mobile-app-based learning group against a physical-card learning group: the app group's scores rose from a pre-test mean of 5.92 to a post-test mean of 8.56, while a separate group's difference was smaller — illustrating exactly the difference-of-differences logic in Section 4's table, used specifically to determine which of two training methods (Topic 24) produced the larger genuine gain, not just which group scored higher overall.

๐ŸŒพ Extension Angle

A published study evaluating WhatsApp-delivered messages on sugarcane cultivation practices used a "before-after without control group" design with paired analysis — a design choice directly shaped by practical field constraints, since maintaining a genuinely separate control group of farmers is often logistically difficult in real extension settings, even though it would strengthen the evaluation's rigor per Topic 26's accuracy standard.

๐Ÿ‡ฎ๐Ÿ‡ณ Indian Institutional Context

These exact tools — mean, paired t-test, correlation — are the standard statistical toolkit in evaluation studies published in ICAR's own Indian Journal of Extension Education, meaning this topic's content reflects live, current Indian extension research practice rather than an imported or generic statistics curriculum.

๐Ÿ”„ Beyond Agriculture

The same toolkit is used to evaluate capacity development programmes for India's ASHA (Accredited Social Health Activist) and Anganwadi health and nutrition workers — recent studies comparing app-based and physical-card refresher training methods for these community health workers used the identical pre-test/post-test, paired-comparison, and intervention-group-comparison logic covered in Sections 2 and 4 above, showing this statistical toolkit is genuinely sector-agnostic across India's community worker training systems.

Frequently Asked Questions

Why is the paired t-test used so much more often than the independent t-test in extension evaluation? +
Because most extension training evaluations use a before/after design on the same group of trainees, which is exactly what the paired t-test is built for. An independent t-test requires two genuinely separate groups (such as a trained group and an untrained comparison group), which is often harder to arrange in real field settings, as Section 5's WhatsApp study example shows.
Does a significant paired t-test result prove the training caused the knowledge gain? +
It shows the pre/post difference is unlikely to be due to chance, but without a control group, other factors occurring during the same period could also explain some of the gain. This is precisely why a difference-of-differences design comparing two groups, as in the Anganwadi Worker example, offers stronger evidence than a single group's before/after comparison alone.
When should I use a chi-square test instead of a t-test? +
Use a chi-square test when your outcome is categorical — such as whether a farmer adopted a practice or not — rather than a continuous score like a knowledge test result. T-tests compare means of continuous data; chi-square tests compare frequencies or proportions across categories.
๐Ÿ“š References
  • Studies published in the Indian Journal of Extension Education using mean, paired t-test, and correlation for pre-test/post-test knowledge-gain evaluation (e.g., studies on video and WhatsApp-based extension delivery for sugarcane cultivation practices).
  • Jayaratne, K. S. U., Chaudhary, A., & Diaz, J. M. Knowledge testing options in pre-test post-test evaluation design: Implications for Extension program evaluation. Advancements in Agricultural Development.
  • Studies on app-based and card-based refresher training for Anganwadi and ASHA community health workers in India, comparing pre-test/post-test gains across intervention groups.
  • Raab, R. T., Swanson, B. E., Wentling, T. L., & Clark, C. D. (Eds.). (1987). A Trainer's Guide to Evaluation. Rome: Food and Agriculture Organization of the United Nations. [Cross-referenced — see Topics 25, 27, 28.]

Featured Post

Research & Study Toolkit

๐Ÿ”Š Listen to This Page Note: You can click the respective Play button for either Hindi or English below. ...

Research & Academic Toolkit

Welcome to Your Essential Research & Study Toolkit by Dr. Singh—a space created with students, researchers, and academicians in mind. Here you'll find simple explanations of complex topics, from academic activities to ANOVA and reliability analysis, along with practical guides that make learning less overwhelming. To save your time, the site also offers handy tools like citation generators, research calculators, and file converters—everything you need to make academic work smoother and stress-free.

Read the full story →