In many research projects, you want to find out whether and how strongly two variables are related – for example, height ↔ weight, study time ↔ exam grade, or CO₂ emissions ↔ income. Correlation analyses provide simple but important tools for this.
Pearson correlation (r)
Pearson’s product–moment correlationr measures the strength and direction of a linear relationship between two metric variables. Values range from -1 (perfectly negative), through 0 (no linear relationship), to +1 (perfectly positive).
The assumptions for conducting a Pearson correlation are:
- Metric scale (interval/ratio).
- Linear relationship – always check this using a scatterplot!
- No major outliers; they can severely distort r.
- Homoscedasticity and an approximately normal distribution are helpful, but not essential for simply estimating r; they are important for significance tests, however.
This is the formula (but you don’t necessarily need to memorize it):
$r=\frac{\sum_{i=1}^{n}(x_i-\bar x)\,(y_i-\bar y)}{\sqrt{\sum_{i=1}^{n}(x_i-\bar )^2}\,\sqrt{\sum_{i=1}^{n}(y_i-\bar y)^2}}$
In R, you can implement this as follows:
# Beispiel: Körpergröße vs. Gewicht
data <- data.frame(
groesse = c(168,172,181,176,169,185,160,178,190,174),
gewicht = c(62, 70, 88, 75, 65, 92, 55, 80, 95, 72)
)
# Visualisierung
plot(data$groesse, data$gewicht,
xlab = "Größe (cm)", ylab = "Gewicht (kg)",
main = "Scatterplot Größe vs. Gewicht")
abline(lm(gewicht ~ groesse, data), lty = 2)
# Pearson-Korrelation + Test
cor(data$groesse, data$gewicht, method = "pearson")
cor.test(data$groesse, data$gewicht, method = "pearson")
The function cor.test() returns a t-value, the 95% confidence interval, and the p-value in addition to r.
For the interpretation, you should focus mainly on these three values:
- Absolute value of r indicates strength:
- 0 – 0.3: weak
- 0.3 – 0.6: moderate
- 0.6 – 1.0: strong
- Sign indicates direction.
- Significance (p-value) indicates whether the observed relationship could have occurred by chance. However: Significant ≠ practically relevant (keyword: effect size).
Beware the causality trap: “Correlation does not imply causation” (see causality)
Rank correlation (Spearman $\rho$)
Spearman’s $\rho$ (rho) measures the monotonic relationship between two variables by transforming the data into ranks and then applying Pearson’s $r$ to these ranks. It is more robust to outliers & non-normality and is suitable for ordinal data.
You use rank correlation (instead of Pearson correlation) when…
- The data are measured on an ordinal scale or show outliers / skewness.
- The relationship is monotonic, but not necessarily linear (e.g., a saturating effect).
In R, you implement this as follows:
# Beispiel: Stress-Rang vs. Prüfungsrang
set.seed(1)
stress <- sample(1:100, 15) # 1 = wenig Stress
note <- sample(1:100, 15) # 1 = beste Note
cor.test(stress, note, method = "spearman",
exact = FALSE) # exact=FALSE für größere n
The result contains $/rho$ and a p-value.
For the interpretation, you could say (assuming $\rho=-0.72 (p < 0.01)$): A higher stress rank (more stress) is monotonically associated with poorer exam ranks.
Kendall’s tau ($\tau$)
Kendall’s $\tau$ is based on pairwise comparisons (concordant vs. discordant pairs). It is particularly robust with small samples and many ties. Values: $[-1,1]$.
cor.test(stress, note, method = "kendall")
When is Kendall’s tau better than Spearman’s?
- Many ties (identical values) in the ranks.
- Very small n (< 20).
Visualization
- Basic plot
plot(x, y); abline(lm(y ~ x)) ggplot2with a smooth linelibrary(ggplot2) ggplot(data, aes(groesse, gewicht)) + geom_point() + geom_smooth(method = "lm", se = FALSE)- Correlation matrix as a heatmap
library(reshape2); library(ggplot2) m <- cor(data, method = "spearman") m2 <- melt(m) ggplot(m2, aes(Var1, Var2, fill = value)) + geom_tile() + scale_fill_gradient2(limits = c(-1,1)) + theme_minimal()
Comparison & decision aid
| Criterion | Pearson $r$ | Spearman $\rho$ | Kendall $\tau$ |
|---|---|---|---|
| Data level | Metric | Ordinal+ | Ordinal+ |
| Relationship | Linear | Monotonic | Monotonic |
| Robust to outliers | No | Moderate | High |
| Small n + ties | Moderate | Moderate | Good |
| Test statistic | t-distribution | t-approximation | z-approximation |
Example study
You collect data from 60 students: mg of caffeine per day vs. reactiontime (ms).
set.seed(42)
koffein <- rpois(60, lambda = 200)
reaktion <- 300 - 0.15*koffein + rnorm(60, 0, 20)
df <- data.frame(koffein, reaktion)
Then let’s take a look at the whole thing visually and analytically:
plot(df$koffein, df$reaktion) # Streuplot
cor.test(df$koffein, df$reaktion)
Result: $r\approx -0.65, p < 0.001$ → higher caffeine consumption is (linearly) associated with shorter reaction time.
Because outliers (> 450 mg) appear, you should also compare it with Spearman:
cor.test(df$koffein, df$reaktion, method="spearman")
Often, $\rho$ is slightly smaller (e.g., −0.60), but it confirms the monotonic trend.
Alles klar?
Ich hoffe, der Beitrag war für dich soweit verständlich. Wenn du weitere Fragen hast, nutze bitte hier die Möglichkeit, eine Frage an mich zu stellen!
