Measures of central tendency in descriptive statistics
In descriptive statistics, there are a variety of measures of central tendency that help us describe the center of a data distribution—content that should not be missing from any introduction to statistics. These measures of central tendency are essential for understanding the characteristics of a data distribution. This is particularly important when it comes to the use of statistics in psychology and occupational psychology. The most important ones, which we will look at today, are the arithmetic mean, the median, and the mode.
Before you get started, make sure you know what a statistical data table looks like.
Measures of central tendency
Arithmetic mean
The arithmetic mean is the best-known measure of central tendency and is often simply referred to as the average. You calculate it by adding up all the observed values and dividing the result by the number of observations. Here is a simple example:
Imagine that you have the following five observations: 4, 6, 8, 10, 12
You calculate the arithmetic mean (x_bar) as follows:
x_bar = (x1 + x2 + x3 + x4 + x5)/n =
= (4 + 6 + 8 + 10 + 12)/5 =
= 40/5 =
= 8
The arithmetic mean of these data is therefore 8. The arithmetic mean is particularly suitable for symmetric distributions without outliers.
Example in R
You can easily calculate the arithmetic mean in R:
# Example: Calculating the arithmetic mean
values <- c(4, 6, 8, 10, 12)
mean(values)
The result is 8.
Median
The median divides a sorted dataset into two equally large halves. 50% of the values lie below the median, and 50% lie above it. This is particularly useful for asymmetric distributions or outliers.
Example:
If you have the following values:
1, 3, 3, 6, 7, 8, 9
the median is 6, as it lies in the middle. This is especially easy to see here because the values are already sorted in ascending order.
With an even number of data points, you take the average of the two middle values:
Median = (3 + 4)/2 = 3.5
Example in R
You can also easily calculate the median in R:
# Example: Calculating the Median
values<- c(1, 3, 3, 6, 7, 8, 9)
median(values)
The result is 6.
Mode
The mode is the value that occurs most frequently. The mode is particularly useful for nominal data.
Example:
In the data 1, 2, 2, 3, 4, 4, 4, 5, 6, the mode is 4 because this value occurs most frequently.
Example in R
R does not have a built-in function for the mode, but you can easily write your own function:
# Example: Calculating the mode
mode <- function(x) {
uniqx <- unique(x)
uniqx[which.max(tabulate(match(x, uniqx)))]
}
values<- c(1, 2, 2, 3, 4, 4, 4, 5, 6)
mode(values)
The result is 4.
Arithmetic Mean vs. Median vs. Mode
It is important to know when each measure of central tendency is most appropriate. The arithmetic mean is useful for symmetrical distributions without outliers. The median is more robust to outliers because it depends only on the position of the values. The mode is particularly useful for nominal data.
Rules of Central Tendency
Measures of central tendency can also help assess the shape of a distribution:
- Symmetrical distribution: The arithmetic mean, median, and mode are equal.
- Left-skewed distribution: The arithmetic mean is greater than the median, which in turn is greater than the mode.
- Right-skewed distribution: The arithmetic mean is smaller than the median, which in turn is smaller than the mode.
Conclusion
Measures of central tendency are essential tools for describing the center and distribution of data. Depending on how your data are distributed, you choose the arithmetic mean, the median, or the mode. They help you conduct well-founded statistical analyses.
