Simple and Multiple Regression in R

Imagine you want to estimate your exam grade based on a single number, e.g. study time in hours. Or perhaps you have several characteristics of a car (horsepower, weight, cylinders) and want to predict its fuel consumption. In both cases, you are looking for a mathematical model that describes a measurable relationship quantitatively and is useful for making predictions. Welcome to linear regression!

Simple Regression

Simple linear regression models the relationship between a metric outcome variable ($Y$) and exactly one predictor ($X$) using the equation $\[ \hat Y \;=\; b_0 \;+\; b_1\,X, \]$ where $\(b_0\)$ is the intercept and $\(b_1\)$ is the slope (the change in $\(Y\)$ per unit of $\(X\)$). Both parameters are estimated so that the sum of the squared distances (residuals) between the observed and predicted $\(Y\)$ is minimized (the so-called least squares method).

The assumptions of regression are:

  1. Linearity – the mean of YYY changes proportionally to XXX.
  2. Independence of the observations.
  3. Homoscedasticity – the variance of the residuals is constant across all XXX values.
  4. Normal distribution of the residuals.

Here is an example in R:

set.seed(42)
n <- 100
hours <- runif(n, 0, 10)
grade <- 1 + 0.4*hours + rnorm(n, sd = 0.8) # 1 = best grade
exam <- data.frame(hours, grade)
model <- lm(grade ~ hours, data = exam)
summary(model)

summary() returns, among other things:

  • Intercept $b_0$​ – expected grade at 0 hours of study.
  • Slope $b_1$​ – grade improvement per hour of study. A negative value indicates that more studying reduces the grade (on a grading scale where 1 = best).
  • $R^2$ – proportion of the variation in grades explained by study time.

It also makes sense to use visualizations with ggplot2:

library(ggplot2)

ggplot(exam, aes(hours, grade)) +
geom_point() +
geom_smooth(method = "lm", se = FALSE) +
labs(x = "Study time (h)", y = "Grade") +
theme_minimal()

Multiple linear regression

Now we have multiple predictors at the same time: $\hat Y \;=\; b_0 + b_1 X_1 + b_2 X_2 + \dots + b_p X_p$ Multiple regression extends the simple model by adding further predictors. Each coefficient $\(b_j\)$ estimates the effect of $\(X_j\)$, provided that all other variables are held constant (ceteris paribus).

Suppose you want to predict the monthly rent ($Y$) of an apartment based on

  • living area in m² ($X_1$​),
  • year of construction ($X_2$​),
  • distance from the city center in km ($X_3$​)

predict.

set.seed(1)
n <- 120
sqm <- runif(n, 30, 120)
year <- sample(1950:2020, n, replace = TRUE)
dist <- runif(n, 0.5, 15)
rent <- 5*sqm + (2025 - year) * (-1.2) + (-80)*dist + rnorm(n, sd = 150)
flats <- data.frame(rent, sqm, year, dist)

multi <- lm(rent ~ sqm + year + dist, data = flats)
summary(multi)

Interpretation of the excerpts from summary(multi):

MetricMeaning (informal)
Estimateestimated effect (slope)
Std. Erroruncertainty of the estimate
t value, `Pr(>t
Multiple R-squaredproportion of variance explained
Adjusted R-squaredadjusted for the number of predictors

Interactions (moderation

Perhaps floor area has a different effect in outlying areas than in the city center. Then add an interaction:

multi_int <- lm(rent ~ sqm * dist + year, data = flats)  # * = main effects and interaction
anova(multi, multi_int) # F-test: does the interaction improve the model?

An interaction $\((X_1 \times X_2)\)$ allows the effect of $\(X_1\)$ to depend on the value of \(X_2\).

Multicollinearity

Multicollinearity occurs when two or more predictors are strongly correlated. The estimates then become unstable: small changes in the data → large jumps in $\(b_j\)$.

Check using the Variance Inflation Factor (VIF):

library(car)
vif(multi) # Values > 5 (or 10) should be examined

Common pitfalls

PitfallBrief explanationRemedy
ExtrapolationPrediction outside the observed rangeWarning: best avoided
OutliersDistort slopesDiagnose with Cook’s distance; use a robust model if necessary
OverfittingToo many predictorsAdjusted R2R^2R2, cross-validation
HeteroscedasticityUnequal varianceWhite test, robust SE (sandwich)

The R commands at a glance

rCopyEditlm(y ~ x1 + x2, data = df)        # Estimate model
summary(mod)                      # Details
confint(mod)                      # 95% CIs
plot(mod)                         # Basic diagnostics
predict(mod, newdata, interval="prediction")
library(car); vif(mod)            # Multicollinearity
library(ggfortify); autoplot(mod) # ggplot2 diagnostics

Mini-project: mtcars data

data(mtcars)
car_mod <- lm(mpg ~ wt + hp + qsec, data = mtcars)
library(ggfortify)
autoplot(car_mod) # Residuals & QQ plot
  • Research question: How do weight (wt), horsepower (hp), and acceleration (qsec) affect fuel consumption (mpg)?
  • Result: Higher weight & horsepower reduce mpg, while higher qsec (slower acceleration) increase mpg.

Conclusion

With just a few lines of R code and the lm() command, you can build powerful predictive models. The key is not to neglect data quality, checking assumptions, and clear communication. Keep your model as simple as possible, but as complex as necessary.

Alles klar?

Ich hoffe, der Beitrag war für dich soweit verständlich. Wenn du weitere Fragen hast, nutze bitte hier die Möglichkeit, eine Frage an mich zu stellen!

Stelle Dominik eine Frage