Classical Test Theory: An Introduction

Why do we need Classical Test Theory?

Classical Test Theory (CTT) is a fundamental approach in psychology that enables us to measure and evaluate psychological constructs such as intelligence, personality traits, and abilities. These characteristics are often not directly observable and therefore need to be captured through indirect measurements, such as psychometric tests or self-report questionnaires. CTT provides the mathematical foundation for assessing the accuracy and meaningfulness of these measurements.

Example: Imagine that you want to find out how high a student’s mathematical ability is. A single test score may give you a rough idea, but how accurate and reliable is that score really? This is where Classical Test Theory comes in—it helps us understand how to estimate the “true” value of an ability from the observed data.

Basic assumptions of Classical Test Theory

The basic principle of CTT is based on the assumption that every measured quantity, that is, every “observed score” (X), consists of a “true score” (T) and a random measurement error (e):

$$X = T + e$$

Here, T represents the “true” value (true score) that we would like to measure—in other words, the actual level of the ability or trait. The measurement error (error, e) represents random influences that could distort the result. These errors can arise from various factors, such as fatigue, fluctuations in concentration, or inaccurate measuring instruments.

Example: Suppose a student receives a score of 85 on a mathematics test. This is her “observed score.” However, her “true score”—that is, her actual mathematical ability—could be higher or lower, depending, for example, on whether she was tired that day or whether the test took place in a noisy environment.

Important implications that follow from this are:

  • The expected value of the measurement error is $E(e) = 0$
  • There is no relationship between $e$ and $T$; $COV(T, e) = 0$.
  • The measurement errors of multiple tests ($e1, e2$) are independent; $COV(e1, e2) = 0$.

Calculating variance: observed and true scores

CTT assumes that the variance of the observed scores ($Var(X)$) is the sum of the variance of the true scores ($Var(T)$) and the variance of the error ($Var(e)$):

$${Var}(X) = \text{Var}(T) + \text{Var}(e)$$

This assumption helps us quantify the reliability of tests mathematically. In practice, this means that Classical Test Theory allows us to calculate how much of the variance in our test results can be attributed to actual differences in ability and how much to random error.

Reliability: A measure of test consistency

A central concept in CTT is reliability, which indicates how accurately and consistently a test measures. Reliability expresses the ratio of the variance of the true scores to the variance of the observed scores:

$$\text{Rel}(X) = \frac{\text{VAR}(T)}{\text{VAR}(X)}$$

High reliability means that the test score is very close to the true score and is only minimally distorted by measurement error. Reliability always ranges between 0 and 1—the closer it is to 1, the more reliable the test result.

Example in R for calculating reliability:
Imagine you have a test with an observed variance of 120 and an estimated variance of the true score of 100. Reliability can then be calculated in the statistical software R as follows:

true_var <- 100
observed_var <- 120
reliability <- true_var / observed_var
reliability

This result shows what percentage of the variance is actually explained by differences in the true score rather than by random errors.

Error theory: The role of measurement error in CTT

An essential component of CTT is the assumption that measurement errors are independent and random and have no systematic influence on the measurement of the true score. This means that errors should tend to cancel each other out across repeated measurements and should not be correlated with the true scores in any way. For example, if a person takes an intelligence test on two different days, the errors on both days should be uncorrelated with the true score.

Example: If you administer a mathematical ability test to a student twice, any errors caused by fatigue or other interfering factors should occur independently of one another.

This assumption is particularly important because it simplifies statistical calculations and analyses. In reality, however, this assumption can be problematic, as errors in real-world tests are often not completely random.

Methods for measuring reliability

CTT offers various procedures for determining a test’s reliability:

  1. Test-retest reliability: This method measures the stability of a test over time. The same test is administered to a group of people at two different points in time, and the correlation between the two test results is calculated. A high correlation value indicates high reliability.
  2. Internal consistency: This method determines the consistency within a test by calculating the correlation between different parts of the same test. A frequently used formula for this is Cronbach’s alpha, which uses the following formula:

    $$\alpha = \frac{k}{k-1} \left( 1 – \frac{\sum \text{VAR}(e)}{\text{VAR}(X)} \right)$$

    Here, ($k$) is the number of items in the test, and ($\sum \text{VAR}(e)$) is the sum of the item error variances.
  3. Split-half reliability: This method divides the test into two halves and calculates the correlation between the results of both halves. This correlation is then converted into an overall reliability estimate.

Example in R for Cronbach’s alpha:
If you want to calculate internal consistency, you can calculate Cronbach’s alpha in R, for example, using the “psych” package:

# Installation des psych Pakets (falls noch nicht installiert)
# install.packages("psych")

# Bibliothek laden
library(psych)

# Beispiel-Datensatz (Test mit mehreren Items)
scores <- data.frame(
  Item1 = c(5, 4, 4, 5, 5, 4),
  Item2 = c(4, 3, 4, 4, 5, 3),
  Item3 = c(5, 5, 4, 4, 4, 4),
  Item4 = c(3, 4, 5, 4, 3, 4)
)

# Berechnung von Cronbachs Alpha
alpha(scores)

Criticism of Classical Test Theory

Classical Test Theory is a widely used and useful model, but it also has weaknesses. A frequent criticism is the assumption that errors are unsystematic and independent. In reality, tests often show systematic errors that can be influenced by factors such as testing conditions, test motivation, or cultural differences.

Another disadvantage is that CTT assumes that all items in a test contribute equally to measuring the construct, which is not always the case for complex tests. Item response theory (IRT) was developed to address these kinds of problems, as it enables more complex models and analyses. However, IRT is more computationally demanding and requires larger samples.

Applying Classical Test Theory in Practice

In psychology, Classical Test Theory is frequently used to develop and evaluate questionnaires, intelligence tests, and other psychological tests. One example is calculating the reliability of a personality test, where it is important that the results are independent of external influences and fluctuations in testing conditions.

Illustrative Visualization of Test Results

It is often helpful to present test results graphically in order to visualize the dispersion and potential outliers. For example, a simple box plot can show the variance and central values. Here is an example in R:

# Beispielhafte Testergebnisse
test_scores <- c(110, 105, 120, 115, 118, 112)

# Boxplot erstellen
boxplot(test_scores, main="Boxplot der Testwerte", ylab="Testwert")

Conclusion

Classical Test Theory provides a solid foundation for evaluating psychological tests and questionnaires and determining their reliability and accuracy. Despite its assumptions and limitations, it is an indispensable tool in psychological assessment. CTT enables us to develop and apply differentiated yet straightforward procedures for measuring and calculating psychological characteristics. Particularly in practical applications—whether in clinical psychology, occupational psychology, or educational research—CTT remains an essential component of test development and evaluation.