Simple Linear Regression
Simple linear regression is a fundamental statistical method that allows you to examine relationships between two variables and make predictions. Simple linear regression is an important tool, particularly in the social sciences, economics, and many other disciplines, for gaining insights from data. But what is it actually about, and how do you apply this method? Let’s take a detailed look!
What is Simple Linear Regression?
Simple linear regression describes the relationship between an independent variable X (for example, the number of hours you studied) and a dependent variable Y (for example, your test score). The goal is to express this relationship using a straight line. This line has the following form:

Using the least squares method, you determine α and β in order to keep the deviations of the actual values y_i from the predicted values as small as possible.
A Simple Example
Suppose you want to find out whether there is a relationship between living space (in square meters) and net rent (in euros). You collect data and obtain the following table:
| Living Space (sqm) | Net Rent (Euros) |
|---|---|
| 50 | 450 |
| 60 | 500 |
| 70 | 550 |
| 80 | 600 |
| 90 | 650 |
Now you want to model the relationship using linear regression. To do this, you use the method of least squares to determine the values of α and β.
Calculation in R
In R, you can perform this calculation very easily. Here is the code:
# Enter the living area (x) and net rent (y) data
living_area <- c(50, 60, 70, 80, 90)
net_rent <- c(450, 500, 550, 600, 650)
# Perform linear regression
model <- lm(net_rent ~ living_area)
# Display a summary of the results
summary(model)
# Plot the data and the regression line
plot(living_area, net_rent, main = "Linear Regression: Living Area and Net Rent",
xlab = "Living Area (sqm)", ylab = "Net Rent (Euros)", pch = 19)
abline(model, col = "blue")
With this code, you perform a simple linear regression in R and immediately obtain an overview of the estimated values for α and β. You also generate a plot showing the data points and the regression line.
Interpretation of the results
Suppose the regression results give you the following estimates:

This means that the net rent y increases by an average of 5 euros when the living area increases by 1 square meter. The value α = 200 means that the estimated net rent for an apartment with 0 square meters would be 200 euros—which of course makes little practical sense, but serves as a mathematical value here.
Checking the model: residual analysis
An important prerequisite for simple linear regression is that the error terms ε_i are randomly distributed and have constant variance. To check this, you examine the residuals (the difference between the observed and predicted values). In R, you can easily visualize the residuals:
# Plot residuals
plot(residuals(modell), main = "Residual analysis", ylab = "Residuals", xlab = "Index")
abline(h = 0, col = "red")
If the residuals are randomly distributed and show no discernible patterns, your model meets the assumptions.
Conclusion
Simple linear regression is a powerful tool for examining the relationship between two variables. With R, you can quickly create models and check whether the model assumptions are met. In our example, you saw how the living area affects the net rent and how you can represent this relationship with a regression line. Why not try it yourself with your own data?
