Imagine you want to estimate your exam grade based on a single number, e.g. study time in hours. Or perhaps you have several characteristics of a car (horsepower, weight, cylinders) and want to predict its fuel consumption. In both cases, you are looking for a mathematical model that describes a measurable relationship quantitatively and is useful for making predictions. Welcome to linear regression!
Simple Regression
Simple linear regression models the relationship between a metric outcome variable ($Y$) and exactly one predictor ($X$) using the equation $\[ \hat Y \;=\; b_0 \;+\; b_1\,X, \]$ where $\(b_0\)$ is the intercept and $\(b_1\)$ is the slope (the change in $\(Y\)$ per unit of $\(X\)$). Both parameters are estimated so that the sum of the squared distances (residuals) between the observed and predicted $\(Y\)$ is minimized (the so-called least squares method).
The assumptions of regression are:
- Linearity – the mean of YYY changes proportionally to XXX.
- Independence of the observations.
- Homoscedasticity – the variance of the residuals is constant across all XXX values.
- Normal distribution of the residuals.
Here is an example in R:
set.seed(42)
n <- 100
hours <- runif(n, 0, 10)
grade <- 1 + 0.4*hours + rnorm(n, sd = 0.8) # 1 = best grade
exam <- data.frame(hours, grade)
model <- lm(grade ~ hours, data = exam)
summary(model)
summary() returns, among other things:
- Intercept $b_0$ – expected grade at 0 hours of study.
- Slope $b_1$ – grade improvement per hour of study. A negative value indicates that more studying reduces the grade (on a grading scale where 1 = best).
- $R^2$ – proportion of the variation in grades explained by study time.
It also makes sense to use visualizations with ggplot2:
library(ggplot2)
ggplot(exam, aes(hours, grade)) +
geom_point() +
geom_smooth(method = "lm", se = FALSE) +
labs(x = "Study time (h)", y = "Grade") +
theme_minimal()
Multiple linear regression
Now we have multiple predictors at the same time: $\hat Y \;=\; b_0 + b_1 X_1 + b_2 X_2 + \dots + b_p X_p$ Multiple regression extends the simple model by adding further predictors. Each coefficient $\(b_j\)$ estimates the effect of $\(X_j\)$, provided that all other variables are held constant (ceteris paribus).
Suppose you want to predict the monthly rent ($Y$) of an apartment based on
- living area in m² ($X_1$),
- year of construction ($X_2$),
- distance from the city center in km ($X_3$)
predict.
set.seed(1)
n <- 120
sqm <- runif(n, 30, 120)
year <- sample(1950:2020, n, replace = TRUE)
dist <- runif(n, 0.5, 15)
rent <- 5*sqm + (2025 - year) * (-1.2) + (-80)*dist + rnorm(n, sd = 150)
flats <- data.frame(rent, sqm, year, dist)
multi <- lm(rent ~ sqm + year + dist, data = flats)
summary(multi)
Interpretation of the excerpts from summary(multi):
| Metric | Meaning (informal) |
|---|---|
Estimate | estimated effect (slope) |
Std. Error | uncertainty of the estimate |
t value, `Pr(> | t |
Multiple R-squared | proportion of variance explained |
Adjusted R-squared | adjusted for the number of predictors |
Interactions (moderation
Perhaps floor area has a different effect in outlying areas than in the city center. Then add an interaction:
multi_int <- lm(rent ~ sqm * dist + year, data = flats) # * = main effects and interaction
anova(multi, multi_int) # F-test: does the interaction improve the model?
An interaction $\((X_1 \times X_2)\)$ allows the effect of $\(X_1\)$ to depend on the value of \(X_2\).
Multicollinearity
Multicollinearity occurs when two or more predictors are strongly correlated. The estimates then become unstable: small changes in the data → large jumps in $\(b_j\)$.
Check using the Variance Inflation Factor (VIF):
library(car)
vif(multi) # Values > 5 (or 10) should be examined
Common pitfalls
| Pitfall | Brief explanation | Remedy |
|---|---|---|
| Extrapolation | Prediction outside the observed range | Warning: best avoided |
| Outliers | Distort slopes | Diagnose with Cook’s distance; use a robust model if necessary |
| Overfitting | Too many predictors | Adjusted R2R^2R2, cross-validation |
| Heteroscedasticity | Unequal variance | White test, robust SE (sandwich) |
The R commands at a glance
rCopyEditlm(y ~ x1 + x2, data = df) # Estimate model
summary(mod) # Details
confint(mod) # 95% CIs
plot(mod) # Basic diagnostics
predict(mod, newdata, interval="prediction")
library(car); vif(mod) # Multicollinearity
library(ggfortify); autoplot(mod) # ggplot2 diagnostics
Mini-project: mtcars data
data(mtcars)
car_mod <- lm(mpg ~ wt + hp + qsec, data = mtcars)
library(ggfortify)
autoplot(car_mod) # Residuals & QQ plot
- Research question: How do weight (
wt), horsepower (hp), and acceleration (qsec) affect fuel consumption (mpg)? - Result: Higher weight & horsepower reduce mpg, while higher
qsec(slower acceleration) increase mpg.
Conclusion
With just a few lines of R code and the lm() command, you can build powerful predictive models. The key is not to neglect data quality, checking assumptions, and clear communication. Keep your model as simple as possible, but as complex as necessary.
Alles klar?
Ich hoffe, der Beitrag war für dich soweit verständlich. Wenn du weitere Fragen hast, nutze bitte hier die Möglichkeit, eine Frage an mich zu stellen!
