Introduction to Item Response Theory (IRT)
Item Response Theory (IRT) is a modern method for developing and analyzing psychometric tests, which is widely used particularly in psychology and the social sciences. Unlike Classical Test Theory (CTT), which makes assumptions about test scores and measurement errors, IRT focuses on the relationship between test responses and the underlying traits of the person being assessed.
What is Item Response Theory?
Item Response Theory (IRT), also known as probabilistic test theory, describes the probability that a person will answer a particular test item correctly. This probability depends on two central factors: the person’s ability and the difficulty of the item. This theory provides a more nuanced perspective on test responses and helps make tests more precise.
Definition: Item Response Theory (IRT): A psychometric model that examines the relationship between responses to test items and the underlying abilities or traits of the people being assessed.
Measurement in IRT vs. Classical Test Theory (CTT)
In classical test theory (CTT), an observed score is assumed to be a mixture of a “true” score and measurement error. The aim is to minimize the proportion of error in order to achieve higher reliability. However, CTT assumes that all items are equivalent, which can mean that differences between individual test-takers are not always captured accurately.
Item response theory (IRT), by contrast, uses probabilistic models that take into account differences in the difficulty and discrimination of each item. This means that IRT enables more accurate estimates of the underlying trait (e.g., pain intensity) by incorporating the test-taker’s response pattern and the specific characteristics of each item.
Basic Assumptions of IRT
IRT is based on several key assumptions that distinguish it from classical test theory:
- Items as indicators of latent traits
Each item in a test is considered an indicator of a latent ability or trait. These latent traits cannot be observed directly but can only be measured indirectly through the person’s responses to the items. - Influence of ability and difficulty
The probability that a person answers an item correctly depends both on the person’s ability and on the item’s difficulty. The higher the ability and the lower the difficulty, the more likely a correct answer becomes. - Unidimensionality and local independence
A test is considered unidimensional when all items measure the same latent trait. In addition, it is assumed that responses to the items are independent of one another once the latent trait has been taken into account (local independence).
The Meaning of Person and Item Parameters
In IRT, person parameters and item parameters are central concepts:
- Person parameters ($theta$, theta) represent a person’s ability or trait in relation to the latent construct.
- Item parameters describe specific properties of the items, such as their difficulty ($b$).
The most important ones are:
Item difficulty: Difficulty describes the point on the ability continuum at which a person has a 50% probability of answering an item correctly. This point is referred to as the “median value” and lies where people with an average ability level are just able to provide a correct answer.
Item discrimination: Discrimination indicates how sensitively an item responds to differences in ability levels. An item with high discrimination can detect very fine-grained differences and therefore helps classify people more precisely.
A classic example of calculating item probabilities is the dichotomous Rasch model, which is used when the response options for the items are binary (e.g., correct/incorrect).
Formula in the Rasch model:
The probability $P(X = 1|theta)$ of answering an item correctly is calculated as follows:
$$P(X_{ni} = 1) = \frac{e^{\theta_n – \beta_i}}{1 + e^{\theta_n – \beta_i}}$$
Here:
- $\theta_n$: Ability parameter of respondent $n$
- $\beta_i$: Difficulty parameter of item $i$
Example: Imagine you are developing an intelligence test with several items. One item could, for example, involve solving a mathematical problem. Let’s assume that Lisa, a test taker, has an ability ($theta$) of 1, while the item has a difficulty ($b$) of 0.5. In this case, the probability that Lisa answers the item correctly is higher than 50%, because her ability is higher than the difficulty of the item.
Models of Item Response Theory
Within IRT, there are various models that differ in their complexity and areas of application:
1. Rasch Model
- Dichotomous: The Rasch model is the simplest IRT model and is generally used for items with two response options (e.g., yes/no). It estimates only one parameter: item difficulty.
- The probability of a correct response increases when the test taker’s ability level exceeds the difficulty of the item. The Rasch model is often used in education and psychometrics.
2. Two-Parameter Model
- This model adds a second parameter: discrimination. This makes it possible to assess how well an item distinguishes between people with different ability levels. A high level of discrimination means that the item can detect finer differences along the ability continuum.
3. Three-Parameter Model
- Here, a third parameter, guessing, is introduced. It represents the probability that a person guesses the correct answer. This is particularly relevant for multiple-choice questions and reflects the fact that guessing can influence the accuracy of the measurement results.
Definition: Discrimination parameter: A measure of how well an item differentiates between people with different abilities.
Probability-based approach to IRT
IRT calculates the probability that a person will give a particular response based on their ability level and the characteristics of the item. For example, a person with a high ability level is more likely to be able to solve a difficult item than a person with a lower ability level.
In practice, IRT estimates the probabilities for each response on a continuous spectrum for every item and ability level. This creates a model that shows not only whether an item is solved, but also how the item’s difficulty and discrimination relate to the ability levels of the test takers.
IRT information and reliability
In IRT, the concept of reliability is redefined as information. Information indicates how precisely an item or test represents different ability levels. Information is highest where items best distinguish between different ability levels, and decreases when the ability level differs substantially from the item’s difficulty.
This approach has the advantage that it does not average measurement precision across all ability levels, as is the case in CTT. Instead, IRT can determine precisely at which ability levels the measurement is particularly accurate. Information functions can be created for individual items as well as for the test as a whole to show where the test performs well.
Advantages and applications of IRT
IRT offers numerous advantages over classical test theory:
1. Scale development and evaluation:
IRT is a valuable tool for selecting and developing scales. Analyzing the information function helps ensure that the scales measure precisely, especially in the relevant ability ranges. For example, a scale for measuring clinical depression should perform particularly well in the range from moderate to severe depression.
2. Scale alignment and linking:
IRT can help align different scales that measure the same construct. This is useful in meta-analyses or longitudinal studies in which different instruments were used. By mapping scores from different scales onto a common metric, IRT enables meaningful comparisons between studies.
3. Computer-adaptive tests (CAT):
The CAT approach optimizes test administration by selecting questions that match the test taker’s ability level. Based on the initial responses, the test is adapted so that the items increasingly match the test taker’s abilities. This reduces the number of questions needed for an accurate assessment, thereby reducing the burden of testing and increasing measurement accuracy.
Figures illustrating the item response curve
To display the item response curve for an item, you can use the following R code:
# R code for visualizing the item response function for a Rasch model
theta <- seq(-3, 3, length.out = 100)
b <- 0.5 # Item difficulty
# Calculation of response probabilities
p <- exp(theta - b) / (1 + exp(theta - b))
# Plot
plot(theta, p, type = "l", lwd = 2, ylab = "Response probability",
xlab = expression(theta), main = "Item Response Function (Rasch Model)")
abline(h = 0.5, col = "red", lty = 2)
Practical application of IRT in psychology
An example of the practical application of IRT is the development of personality questionnaires. Such a questionnaire could contain items that measure self-confidence. IRT analysis helps identify the items that best differentiate between people with different levels of self-confidence. The analysis could also show that certain items are particularly well suited to capturing high or low levels of self-confidence.
