Exam questions on the t-test

Task

Case description

The aim of this study was to investigate whether a two-week mindfulness training significantly reduces students’ subjective stress levels. To this end, a sample of N = 20 bachelor’s students was recruited, each of whom rated their stress level before and after the training on a scale from 1 (very low) to 10 (very high). The data were analyzed as a paired sample. A t-test for dependent samples, diagnostic assumptions (normal distribution, equality of variances), and a nonparametric Wilcoxon signed-rank test were conducted. The aim was to critically evaluate both the effectiveness of the training and the suitability of the statistical tests used.

R script

set.seed(123)
stress_vorher <- rnorm(20, mean = 7, sd = 1)
stress_nachher <- stress_vorher - rnorm(20, mean = 0.8, sd = 0.5)
daten <- data.frame(stress_vorher, stress_nachher)

# Data overview
head(daten)

# Checking assumptions
shapiro.test(daten$stress_vorher)
shapiro.test(daten$stress_nachher)
var.test(daten$stress_vorher, daten$stress_nachher)

# T-test for dependent samples
t.test(daten$stress_vorher, daten$stress_nachher, paired = TRUE)

# Wilcoxon test
wilcox.test(daten$stress_vorher, daten$stress_nachher, paired = TRUE)

Output

Data overview:

stress_beforestress_after
7.4395246.937332
6.7698235.777982
7.5587086.764638
8.0705086.985201
6.1292885.318256
6.7150655.758218

Shapiro–Wilk test for normality:

  • Before: W = 0.969, p = 0.681
  • After: W = 0.975, p = 0.787

F-test for homogeneity of variances:

  • F = 0.963, p = 0.930

Paired-samples t-test:

  • t = 5.201, df = 19, p < 0.001
  • 95% confidence interval: [0.662, 1.391]
  • Mean difference: 1.027

Wilcoxon signed-rank test:

  • V = 193, p < 0.001

Questions

Interpreting the output

  1. What does the t-test show regarding the effect of the mindfulness training?
  2. How do you interpret the 95% confidence interval of the t-test?
  3. What conclusion does the Wilcoxon test allow you to draw compared with the t-test?

Theoretical and methodological foundations

  1. State the null and alternative hypotheses of the paired-samples t-test in this case.
  2. What assumptions must be met for the paired-samples t-test to provide valid results?
  3. What role does testing for homogeneity of variances play in paired t-tests?

Reflection and application

  1. Suppose the Shapiro–Wilk tests had produced significant p-values—how would you have had to respond methodologically?
  2. Discuss this example to determine when it makes sense to perform the Wilcoxon test in addition to the t-test.

Model answer

Interpretation of the output

What does the t-test show regarding the effect of the mindfulness training?

The paired-samples t-test shows a significant difference between the stress levels before and after the mindfulness training (t = 5.201, p < 0.001). This suggests that the training led to a significant reduction in participants’ subjective stress.

How do you interpret the 95% confidence interval of the t-test?

The confidence interval [0.662, 1.391] indicates that, with 95% confidence, the true mean reduction in stress in the population lies between 0.662 and 1.391 points on the scale. Since the interval does not contain zero, it supports the significance of the effect.

What conclusion does the Wilcoxon test allow in comparison with the t-test?

The Wilcoxon signed-rank test also confirms a significant difference (V = 193, p < 0.001). As a nonparametric procedure, it does not depend on the assumption of normality and serves as a robust complement to the t-test. The result therefore strengthens the evidential basis of the analysis.

Theoretical and methodological foundations

Null and alternative hypotheses of the t-test:

  • H0: There is no difference in the average stress level before and after the training (mean difference = 0).
  • H1: There is a difference in the average stress level before and after the training (mean difference ≠ 0).

Assumptions for the t-test with paired samples:

  1. Differences between measurement time points are normally distributed (tested using the Shapiro–Wilk test).
  2. The measurement values are measured at the interval level.
  3. No outliers causing substantial distortion of the distribution.

Role of homogeneity of variance:

Homogeneity of variance is not required for paired t-tests because the differences are analyzed. Testing the variance is therefore not methodologically necessary, although it is sometimes still carried out.

Reflection and application

What to do if normality is significantly violated?

If the Shapiro-Wilk tests had been significant, this would indicate a violation of the assumption of normality. In this case, a nonparametric test such as the Wilcoxon test would be appropriate because it does not make any distributional assumptions.

When is the Wilcoxon test useful?

The Wilcoxon test is useful:

  • for small samples (such as here, N = 20), where normality is difficult to assess,
  • when there are noticeable outliers or skewness in the differences,
  • to validate the results of the t-test (robust sensitivity analysis).