Extension of Classical Test Theory

Classical Test Theory (CTT) is a fundamental model for assessing the accuracy and validity of tests, but it has its limitations. In this article, we take a detailed look at some advanced concepts in CTT as well as criticisms that have led to new approaches in test theory.

Classical Test Theory (CTT): CTT assumes that every measured test performance consists of a “true” value and a random error value. This theory is often used to determine how accurately a test measures a particular psychological or cognitive aspect. You can read more about it here.

Correction for attenuation (attenuation correction)

Attenuation correction may sound complicated at first, but it is essentially a method for obtaining more accurate results in statistics. Imagine that you are taking a test—say, a math test—and you want to know how well the test really measures someone’s mathematical ability.

What is the problem?

Every test has some inaccuracies—you may have had a bad day, or the test may contain questions that are not entirely clear. These small errors mean that the test result does not reflect with 100% accuracy how good someone really is at math. In statistics, this accuracy is called reliability.

Why is attenuation correction important?

Let’s assume you want to know how strongly math grades are related to physics grades (in other words, how similar the grades in both subjects are). But if the math test is not completely accurate, the correlation—the relationship between the two grades—will appear weaker than it actually is. This is where the attenuation correction comes in.

What does the attenuation correction do?

The attenuation correction “corrects for measurement error.” It attempts to reduce the impact of these inaccuracies and gives you a better estimate of how strongly math and physics are actually related. You could think of it as showing you what the correlation would be if the test were perfect—completely free of inaccuracies.

Formula

The formula for the attenuation correction looks like this:

$$
r_{xy_{\text{kor}}} = \frac{r_{xy}}{\sqrt{r_{xx} \cdot r_{yy}}}
$$

Here is what the individual symbols mean:

  • $r_{xy_{\text{kor}}}$: This is the result, the “corrected” correlation. It shows how strong the relationship between two variables (e.g., math and physics) would be if both tests were perfectly reliable, meaning free of error.
  • $r_{xy}$​: This is the “uncorrected” correlation, meaning the original correlation we calculated before applying the attenuation correction.
  • $r_{xx}$ and $r_{yy}$​: These are the reliabilities of the two tests (e.g., how accurately the math test and the physics test measure). They indicate how reliable or “error-free” each test is on its own. Reliability always lies between 0 and 1, where 1 means that the test is perfect and 0 means that the test is completely unreliable.

Suppose you have an uncorrected correlation $r_{xy}$​ of 0.5 between mathematics and physics, but the reliability for the mathematics test ($r_{xx}$​) is 0.8 and for the physics test ($r_{yy}$​) is also 0.8.

The calculation then looks like this:

$$r_{xy_{\text{corr}}} = \frac{0.5}{\sqrt{0.8 \cdot 0.8}} = \frac{0.5}{\sqrt{0.64}} = \frac{0.5}{0.8} = 0.6255$$

The corrected correlation would therefore be 0.625 instead of 0.5.

The influence of test length on reliability

Another aspect is the insight that a test’s reliability can be increased by making the test longer. The longer a test, the more data points are collected and the more reliable the result becomes. This is because doubling the test length actually leads to a doubling of the error variance, but to a quadrupling of the true variance. Here is the explanation:

In Classical Test Theory, a test result (that is, the observed score) consists of two components:

  1. True score: The actual score that reflects the test participant’s ability or trait.
  2. Error score: Random deviations caused by various factors (e.g., fatigue, inattention).

Each component has its own variance:

  • The true variance is the variance of the true scores across all test participants.
  • The error variance is the variance of the error scores, that is, the random fluctuation that makes the observed score less precise.

When the length of a test is doubled, the following happens:

  1. Error variance: Since doubling the test also doubles the number of items, the number of random sources of error (i.e., error components) doubles as well. Therefore, the error variance doubles.
  2. True variance: The true variance, however, grows proportionally to the number of items. Since the test is doubled in length, the true variance increases by the square of the lengthening factor. In this case, this means that the true variance quadruples (i.e., ( 2^2 = 4 )), because the true variance is estimated more stably over the longer measurement.

Let’s look at this again mathematically:

Let’s assume:

  • The true variance of a test is ( $\sigma^2_{\text{true}}$ ).
  • The error variance is ( $\sigma^2_{\text{error}}$ ).

If the number of items is doubled:

  • The error variance becomes ( $2 \cdot \sigma^2_{\text{error}}$ ), because the sources of error occur twice as often.
  • The true variance becomes ( $4 \cdot \sigma^2_{\text{true}}$ ), since the true information is consistent in each item and the information becomes more stable through the longer test.

Since the reliability of a test is defined as the proportion of true variance to total variance, a greater true variance (relative to the error variance) means higher reliability. Therefore, doubling the test length makes the test more reliable, because the true variance grows faster than the error variance.

We can also use the Spearman–Brown prophecy formula for this:

$$ \text{Reliability}{\text{new}} = \frac{k \cdot \text{Reliability}{\text{old}}}{1 + (k – 1) \cdot \text{Reliability}_{\text{old}}} $$

Here, ( k ) represents the test-lengthening factor (e.g., doubling the test length means ( k = 2 )).

Criticisms of Classical Test Theory

Despite its widespread use, Classical Test Theory (CTT) has some fundamental limitations.

Here is a more detailed discussion of these three criticisms of Classical Test Theory (CTT):

Static assumptions about “true” values

Classical Test Theory (CTT) assumes that the true value of a particular trait remains stable for a person over short periods of time. This means that, for example, a person would have almost identical “true” values on an intelligence test today and tomorrow, regardless of possible fluctuations caused by their current mood or external influences. This assumption works well for stable personality traits or cognitive abilities that generally remain constant over shorter periods of time.

However, many psychological characteristics are not constant. Moods, for example, can change over the course of a day—a test taken in the morning could produce completely different results from the same test taken in the evening. Likewise, cognitive performance can fluctuate considerably due to sleep deprivation, stress, or other short-term influences. However, because CTT assumes a constant true value, it overlooks such fluctuations and therefore cannot make reliable statements about changing states.

Example: Imagine that a person takes a concentration test on a stressful Monday morning and then takes it again on Friday evening when they are relaxed. This person’s true concentration ability could differ on the two days, but CTT cannot capture this because it assumes that true concentration is always the same.

Error scores are random and uncorrelated

CTT continues to assume that error scores (that is, deviations from the true score) are random and independent across different test items. This means that an error on one question has no influence on the answers to other questions. These error scores should also be evenly distributed in both directions (positive and negative) across the entire group, so that they cancel each other out on average.

In practice, however, there are numerous influences that violate this assumption. Systematic factors such as test anxiety, physical discomfort, or persistent distractions could affect several test questions and thereby lead to correlated errors. For example, if a person performs worse on the first questions because of nervousness, this could also impair their performance on subsequent questions. This undermines the assumption that error scores are random and independent, which can lead to biased test results.

Example: A student might work more slowly on the first questions because of test anxiety, which makes them even more nervous and causes them to perform worse on the following questions as well. However, this relationship between errors on different questions is not taken into account in CTT.

Failure to account for cognitive processes

Classical Test Theory (CTT) focuses exclusively on the formal aspects of test results, namely the results in numerical form, and places little emphasis on the processes that lead to these results. The theory does not consider how test takers arrive at their answers or which cognitive strategies and processes underlie their responses. Instead, CTT views each response as a simple expression of the true score plus a random error component.

Modern models such as Item Response Theory (IRT) go one step further by attempting to incorporate the cognitive processes underlying responses. For example, IRT takes into account that different questions have different levels of difficulty and that people respond differently to these questions. IRT could therefore explain why a person with average ability finds it easier to answer simpler questions but experiences greater difficulties as the questions become more challenging. CTT, by contrast, assumes that all questions in a test make the same contribution to measurement.

Example: Let us assume that two people achieve the same score on a knowledge quiz. CTT would conclude that both people have the same level of knowledge. IRT, however, could show that one person answered difficult questions correctly but made mistakes on easier questions, while the other person got only the easy questions right. IRT would therefore arrive at a more differentiated assessment.

Alles klar?

Ich hoffe, der Beitrag war für dich soweit verständlich. Wenn du weitere Fragen hast, nutze bitte hier die Möglichkeit, eine Frage an mich zu stellen!

Stelle Dominik eine Frage