Measures of association for discrete variables
In statistics, we often encounter discrete (categorical) variables, such as gender, level of education, or response categories in a survey. To analyze the relationship between such variables, we use measures of association specifically designed for discrete data. In this blog post, you will learn which measures are available, how they work, and how to calculate them in R.
Contingency Tables: The Starting Point
The most common approach for representing the relationship between discrete variables is the contingency table. This table shows how often each combination of categories of the two variables occurs in your data. A simple example could be an investigation into whether gender (male/female) influences the choice of degree program (e.g., mathematics/computer science).
Example of a 2×2 contingency table:
| Mathematics | Computer Science | |
|---|---|---|
| Male | 40 | 30 |
| Female | 25 | 35 |
The question now is whether there is a statistically significant relationship between gender and the choice of degree program.
Chi-Square Test: A First Tool
The chi-square test is the standard procedure for checking whether a relationship exists between the variables. The test compares the observed frequencies in the contingency table with the expected frequencies if the variables were independent.
The formula for the chi-square value is:
$\chi^2 = \sum \frac{(O_i – E_i)^2}{E_i}$
Here, $O_i$ represents the observed frequencies and $E_i$ the expected frequencies.
Calculation in R
Let’s take a look at how you can calculate this test in R. Suppose you have the table shown above and want to perform the chi-square test.
# Enter contingency table
daten <- matrix(c(40, 30, 25, 35), nrow = 2, byrow = TRUE)
colnames(daten) <- c("Mathematics", "Computer Science")
rownames(daten) <- c("Male", "Female")
# Perform chi-square test
chisq.test(daten)
This code gives you the result of the chi-square test and shows whether the association is statistically significant. If the p-value is less than 0.05, you can assume that there is an association between the two variables.
Odds ratio: A useful measure for 2×2 tables
For 2×2 contingency tables, you can also use the odds ratio to quantify the strength of the association. The odds ratio (OR) calculates the ratio of the odds of a particular event occurring in one group compared with another group.
The formula for the odds ratio in a 2×2 table is:
$OR = \frac{(a/c)}{(b/d)}$
where $a$, $b$, $c$, and $d$ represent the cells of the table.
Calculating the odds ratio in R:
# Install package for odds ratios
install.packages("epitools")
library(epitools)
# Calculate odds ratio
oddsratio(daten)
Contingency coefficient: An alternative measure
Another measure you can use is the contingency coefficient KKK, which is based on the chi-square value and quantifies the association. The contingency coefficient is particularly useful when you have tables larger than 2×2.
The formula for the contingency coefficient is:
$K = \sqrt{\frac{\chi^2}{\chi^2 + n}}$
Here, $n$ is the total number of observations.
Calculation in R
You can calculate the contingency coefficient with a simple R script:
# Run the chi-square test again
chi_result <- chisq.test(daten)
# Calculate the contingency coefficient
K <- sqrt(chi_result$statistic / (chi_result$statistic + sum(daten)))
K
Conclusion
Measures of association such as the chi-square test, the odds ratio, and the contingency coefficient are indispensable tools for analyzing relationships between discrete variables. They help you interpret data in a clear and concise way and make decisions based on statistical tests.
With R, you can easily calculate these measures and take your data analysis to the next level. Try it yourself and find out which variables are related in your data!
