Item analysis: How to develop high-quality test questions

The quality of a test depends largely on the individual test questions, known as items. Item analysis helps assess and improve the quality of these questions and ensure that they appropriately differentiate among the target group.

What is item analysis?

Item analysis is a statistical procedure for evaluating the quality of test items. The aim is to determine how well an item can differentiate between the performance or characteristics of members of a target group. It is particularly important in the context of questionnaire design is of great importance.

Definition:
Item analysis is a statistical procedure for evaluating the quality of test items. The aim is to determine how well an item can differentiate between the performance or characteristics of members of a target group.

Why is item analysis important?

A good item analysis ensures that a developed test is:

  1. Valid —that is, it measures what it is intended to measure.
  2. Reliable—it produces stable results.
  3. Fair —it is understandable and accessible to all participants.

Imagine you are developing an intelligence test. If some questions are so difficult that almost nobody can answer them, or so easy that everyone answers them correctly, these items provide no useful information. Item analysis helps you identify such weaknesses.

However, keep in mind that item analysis, which takes place after an initial round of data collection, does not replace good planning !

The steps of item analysis

Item analysis can be broken down into three broad steps:

  1. Create a data matrix
  2. Analyze items
    • Analyze difficulty
    • Analyze variance
    • Analyze discrimination
  3. Select items

Step 1: Creating the data matrix

Before the actual analysis begins, the test results are organized in a data matrix (data table). Each row corresponds to a test participant, and each column to an item.


Example:

ParticipantItem 1Item 2Item 3
1101
2011

This standardized structure should always be followed and makes all subsequent steps easier, regardless of the statistical software you choose.

Step 2: Item analysis

For item analysis, basic knowledge of statistics is essential.

Difficulty analysis

Items should be neither too easy nor too difficult, because both would prevent us from distinguishing between cases (this is also referred to as “discriminating”).

The difficulty of an item is calculated as the proportion of test takers who answer it correctly. This value is referred to as the difficulty index ($P_i$).

$P_i = \frac{\text{Number of correct answers to the item}}{\text{Total number of participants}}$

Example:
In a test with 100 participants, 60 people answer a question correctly. The difficulty index is:
$P_i = \frac{60}{100} = 0.6$
The item has a medium level of difficulty.

An ideal difficulty value is often between 0.2 and 0.8.

Note: Here, we are working within the framework of Classical Test Theory (CTT). Item Response Theory (IRT) takes a different view of item difficulty from the one described here.

Difficulty in performance tests

In performance tests, we can explicitly distinguish between correct (R), incorrect (F), and omitted answers (A). These codes (R, F, A) can be entered directly into the data table.

The difficulty is then calculated as follows:

$P_i = \frac{\text{Number of R}}{\text{Total number of participants}}$

Difficulty with personality tests

With personality tests, the distinction between “right” and “wrong” does not seem entirely appropriate. However, because a scale is polarized, high agreement is always indicative of a high level of the trait—so the same logic can be applied.

The only new aspect we need to consider is the possibility of having more diverse response categories. These must be included in the calculation:

$P_i = \frac{\sum_{v=1}^{n} y_{vi}}{n \cdot (k)} \cdot 100$

Here, we use as the basis the maximum possible column total, which is determined by the number of cases $n$ and the number of response categories $k$.

Example: What is the difficulty of an item (response options 1–5) that was answered by three people with 2, 3, and 4?

Substituting these values into the formula gives:

$P_i = \frac{9}{3*5} \cdot 100 = 60%$

Difficulty analysis in R
# Beispiel: Schwierigkeitsanalyse
responses <- c(1, 0, 1, 1, 0, 1, 0, 0, 1, 0)
P_i <- mean(responses)
P_i

Item variance

Item variance indicates how much the responses to an item vary. An item with high variance can better differentiate between people. You can find out how it is generally calculated in the discussion of measures of dispersion.

But this can also be done differently here. We have already looked at item difficulty, which in a way represents the mean. If we start from the example above, for instance, we could calculate: $P_i \cdot k = 0.6 * 5 = 3$. This corresponds to the arithmetic mean. We could therefore also substitute this into the variance formula and rearrange it, which simplifies, for dichotomous items, to the product of the probability $P_i$ and the complementary probability $1-P_i$.

Example:
For an item with $P_i = 0.6$:
Variance=0.6⋅(1−0.6)=0.24\text{Variance} = 0.6 \cdot (1 – 0.6) = 0.24


Item discrimination index

Item discrimination ($r_{it}$) measures how well an item distinguishes between high- and low-performing people.
Item discrimination is the correlation between the item scores and the total test score. Values between 0.4 and 0.7 are considered good. If the correlation is close to 0, this means that the measurement is independent of the total score. A negative correlation coefficient even indicates an inverse relationship—perhaps the item was not reverse-coded correctly?

Example:
An intelligence test contains a question that is answered correctly by almost everyone with a high total score, but by hardly anyone with a low score. This item has high discrimination.

To calculate item discrimination, you need the total test score. This is the arithmetic mean of all items that can be assigned to this test (i.e. the row mean!). Note: The discrimination therefore changes depending on which items are actually selected in the next step! This process needs to be carried out iteratively.

Item discrimination analysis in R
# Beispiel: Trennschärfe
responses <- data.frame(
  Item1 = c(1, 0, 1, 1, 0, 1, 0, 0, 1, 0),
  TotalScore = c(8, 6, 7, 9, 5, 8, 6, 4, 9, 5)
)
cor(responses$Item1, responses$TotalScore)

Step 3: Item selection

Based on the difficulty indices, variances, and discrimination indices, items are selected or removed.

First, we can look at item difficulty. The selection depends primarily on the range across which we want to discriminate reliably. If, for example, the focus is more on the middle values of the scale, item difficulties around 0.50 are particularly interesting. In general, the closer the difficulty is to 0 or 1, the less the item discriminates between cases.

Second, item variance can also be taken into account. Here, items with high variance are given preference,

And finally, we can also consider the discrimination indices of the items.

But be careful: Ultimately, these are purely technical characteristics. They must always be evaluated against the underlying theoretical assumptions.

Alles klar?

Ich hoffe, der Beitrag war für dich soweit verständlich. Wenn du weitere Fragen hast, nutze bitte hier die Möglichkeit, eine Frage an mich zu stellen!

Stelle Dominik eine Frage