Validity in Psychology: Understanding Test Validity
Imagine you’re taking a test—let’s say, an intelligence test. What if the result doesn’t actually measure your intelligence, but rather your mood or concentration on that particular day? This is exactly where validity comes in. It answers the question: “Does the test actually measure what it is supposed to measure?”

What does validity mean?
Definition:
Validity refers to the extent to which a test is valid. It indicates whether interpretations of test scores are appropriate and meaningful. In short: A test is valid if it measures exactly what it is supposed to measure.
However, validity is a complex concept that can be divided into different dimensions. Let’s look at the most important types.
Content validity
This involves examining whether the content of a test comprehensively represents the target characteristic, a latent construct.
Example:
Imagine you are developing a test of mathematical skills for primary school students. If your tasks contain only addition but no multiplication, an essential part is missing. The content validity would then be low. You would not be comprehensively covering the latent construct “mathematical skills.”
注意:
Content validity is often confused with face validity. While content validity is assessed from a subject-matter perspective, face validity is more subjective: Does the test appear meaningful at first glance?
Criterion validity
This type of validity describes how well a test can predict outcomes in relation to external criteria. This is relevant because diagnostic tests are always conducted with a specific goal in mind. In other words, what is it actually about, and what should the instrument provide information for? Criterion validity can be further divided into predictive validity, concurrent validity, and incremental validity.
Predictive validity: For example, an intelligence test predicts later career success. If the predictions are accurate, the test has high predictive validity. This is also referred to as predictive validity. Findings from basic psychological research are, of course, important here—and any biases present in that research must also be critically considered.
Two common forms of bias in criterion validity are:
- Slope bias: The slope of the regression line (the relationship between the test and the criterion) varies between groups. Example: An intelligence test predicts career success less accurately for women than for men. A test would therefore need to be developed that enables equally valid predictions for both groups.
- Intercept bias: The regression constants (the point where the line intersects the Y-axis) differ between groups. Example: Women are predicted to have greater career success than men with the same test score, which can lead to systematic underestimation or overestimation. This could result from gender-specific disadvantages in the work environment.
Convergent validity: A depression questionnaire, for example, produces similar results to a clinical interview. This is also referred to as concurrent validity.
Incremental validity: This concerns the additional knowledge provided by a test. For example, whether a new questionnaire usefully supplements existing diagnostic methods or provides additional information.
Example of incremental validity
The article by Lima et al. (2005; see reference below) examines the incremental validity of the Minnesota Multiphasic Personality Inventory (MMPI-2) in predicting treatment outcomes. A study was conducted with two groups: One group of therapists had access to their patients’ MMPI-2 data, while the other group did not. Treatment outcomes were assessed using criteria such as symptom improvement, the number of sessions, and treatment dropout.
The main findings were:
- Access to MMPI-2 data did not significantly improve treatment outcomes compared with other commonly used diagnostic instruments.
- A significant analysis showed that patients whose therapists had access to the MMPI-2 experienced less improvement in symptoms than those in the control group.
- There were no differences between the groups in the number of sessions or rates of premature treatment termination.
The findings suggest that, in this setting, the MMPI-2 may offer no additional diagnostic or therapeutic value compared with other instruments. Further studies are recommended to identify the conditions under which the MMPI-2 might be more useful.
Reference: Lima, E. N., Stanley, S., Kaboski, B., Reitzel, L. R., Richey, A., Castro, Y., … Joiner, T. E. Jr. (2005). The incremental validity of the MMPI-2: When does therapist access not enhance treatment outcome? Psychological Assessment, 17(4), 462–468. https://doi.org/10.1037/1040-3590.17.4.462
Construct validity
This describes whether a test actually captures a theoretical construct, such as intelligence or personality. This can only be represented in relative terms—and thus the real question is: To what extent can the construct be integrated into the (expected) nomological network using the test? The nomological network comprises variables relevant to the construct, for which relationships are hypothesized.
Smith’s (2005) Review
The article examines the development of the concept of construct validity, introduced by Cronbach and Meehl (1955), and explores its significance for psychological research and clinical assessment.
Main points:
- Definition: Construct validity describes how well a test measures a theoretical, not directly observable construct. This is determined by examining theoretical predictions about relationships with other constructs.
- Challenges: The concept requires continuous evaluation because test results may reflect not only the target construct but also auxiliary theories or methodological weaknesses.
- Five-step model: A systematic approach to construct validation comprises (1) specifying the construct, (2) formulating hypotheses, (3) designing the research, (4) conducting empirical validation, and (5) revising the theory.
- Progress: Theoretical integration and methodological innovations (e.g., multitrait-multimethod analyses) have improved the evaluation and application of construct validity.
- Practical applications: Advances in clinical diagnosis include more differentiated models of personality disorders and the improvement of measurement instruments through rigorous testing and critical review.
Construct validity remains a central component of psychological research and clinical practice. It requires an iterative, critical engagement with theories, methods, and empirical findings in order to develop more precise and reliable measurement instruments.
Reference: Smith, G. T. (2005). On Construct Validity: Issues of Method and Measurement. Psychological Assessment, 17(4), 396–408. https://doi.org/10.1037/1040-3590.17.4.396
Two aspects of construct validity:
- Convergent validity:
Similar tests should produce similar results. Example: Two different IQ tests correlate highly. - Discriminant validity:
Unrelated constructs should not correlate with each other. Example: An intelligence test should not correlate with a mood test.
The multitrait-multimethod matrix
The article by Campbell and Fiske (1959) introduced the concept of the multitrait-multimethod matrix (MTMM) to assess the validity of psychological tests. The authors proposed examining both convergent and discriminant validity in order to comprehensively evaluate a test’s quality.
These are the main points:
- Convergent validity: A test demonstrates convergent validity when measurements of the same construct using different methods correlate highly.
- Discriminant validity: A test demonstrates discriminant validity when it correlates weakly with other constructs, even when similar methods are used.
- MTMM matrix: The matrix combines multiple constructs (traits) and methods to systematically analyze correlations. Four main areas are examined:
- Validity diagonal (high values indicate convergent validity).
- Heterotrait-heteromethod correlations (should be lower than the validity diagonals).
- Heterotrait-monomethod correlations (indicate method effects).
- Patterns of relationships between constructs and methods.
Implications:
- The MTMM matrix helps uncover both method-related biases and weaknesses in construct validity.
- It emphasizes the importance of method independence and calls for multiple procedures to validate measurements.
- The methodology became influential in the development of psychometric instruments and the assessment of validity.
The MTMM matrix is a systematic tool that makes it possible to critically assess the quality and accuracy of psychological tests, and provides a framework for improving measurement instruments by taking trait and method effects into account.
Reference: Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validity by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105.
Challenges in validity
Validity sounds simple in theory, but challenges arise in practice:
- Bias and distortion:
A test can produce different results for different groups.
Example: A job application test unconsciously favors men because it places greater value on typically male communication styles.
Example of gender bias
The article by Boggs et al. (2005; reference below) examines possible gender biases in the DSM-IV diagnostic criteria for four personality disorders: borderline, schizotypal, avoidant, and obsessive-compulsive personality disorder. A sample of 668 clinical patients was examined for functional impairments, analyzing gender differences in the relationship between diagnostic criteria and impairment.
Main findings:
- General results:
- Most diagnostic criteria showed no systematic gender biases.
- Among the criteria that did show bias, this primarily concerned borderline personality disorder (BPD).
- Borderline criteria:
- The criteria “stress-related paranoia,” “affective instability,” “unstable relationships,” and “intense anger” showed some gender-specific differences.
- Women with the same symptoms as men often functioned better overall, indicating a possible underestimation of women’s functional capacity.
- Other disorders:
- Few biases were identified for schizotypal, avoidant, and obsessive-compulsive personality disorders.
- Implications:
- The findings suggest that some BPD criteria may not fully capture the manifestation of the disorder in men.
- Further research is needed to better understand gender differences and adapt the diagnostic criteria.
While most DSM-IV criteria are gender-neutral, there is evidence of diagnostic bias in borderline personality disorder. This highlights the need to carefully review and adapt the criteria to ensure that they are equally valid for men and women.
Reference: Boggs, C. D., Morey, L. C., Skodol, A. E., Shea, M. T., Sanislow, C. A., Grilo, C. M., … Gunderson, J. G. (2005). Differential impairment as an indicator of sex bias in DSM-IV criteria for four personality disorders. Psychological Assessment, 17(4), 492–496. https://doi.org/10.1037/1040-3590.17.4.492
This topic is also particularly important when tests are administered in different languages. On the one hand, a scientific translation process must be followed, but the subsequent rigorous validation of the translated test is equally important.
Example of comparing language versions
The article by Wiebe and Penley (2005; see the reference below) examines the psychometric properties of the Beck Depression Inventory-II (BDI-II) in English and Spanish. The study analyzed the reliability and validity of both language versions using a sample of 895 college students, many of whom were bilingual.
Main findings:
- Reliability: Both versions of the BDI-II showed strong internal consistency (English: Cronbach’s alpha = .89, Spanish: .91) and good test–retest reliability over a one-week period.
- Factor structure: A confirmatory factor analysis showed that the two-factor structure of the English BDI-II (cognitive-affective and somatic symptoms) was also applicable to the Spanish version.
- Cross-language equivalence: Among bilingual participants, there were no significant differences in scores between the two language versions. The order in which the languages were administered had no effect.
- Time effect: Regardless of the language, participants reported fewer depressive symptoms at the second measurement than at the first.
The results show that the Spanish translation of the BDI-II has psychometric properties comparable to those of the English version. This suggests that the instrument can be used reliably in both languages, particularly in nonclinical samples. Further studies are needed to examine its generalizability to clinical populations.
Reference: Wiebe, J. S., & Penley, J. A. (2005). A psychometric comparison of the Beck Depression Inventory-II in English and Spanish. Psychological Assessment, 17(4), 481–485. https://doi.org/10.1037/1040-3590.17.4.481
- Method effects:
Different testing methods can influence the results. The multitrait-multimethod approach helps identify these effects.
Examples from practice
1. Academic performance test:
A teacher develops a test to measure reading comprehension. If the test is based solely on multiple-choice questions, it might measure guessing ability rather than comprehension—a sign of low validity!
2. Job aptitude test:
A company wants to test teamwork skills, but uses a written questionnaire without practical exercises. The result: The test does not really capture how applicants work in a team.
Q&A
Validity describes the accuracy with which a test or measuring instrument measures what it claims to measure. An intelligence test is valid, for example, if it actually captures cognitive ability rather than other factors such as motivation or anxiety. Validity is essential because only valid tests produce scientifically sound and meaningful results that are suitable for answering the underlying research question. Without validity, the results are worthless and can lead to incorrect decisions.
There are three central types of validity: content validity, criterion validity, and construct validity. Content validity examines whether the test content fully covers the target construct. Criterion validity measures the agreement between test results and external criteria—for example, whether an employment test can predict future job performance. Construct validity assesses whether a test actually measures the underlying theoretical construct, such as whether a creativity test truly captures creativity. Each type of validity contributes to the overall validity of the test.
Content validity is assessed through expert judgment. Specialists analyze whether the test items fully and representatively capture the characteristic being measured. For example, in a test of social competence, they might examine whether the items cover all relevant aspects, such as empathy and communication skills. Insufficient content validity could mean that important areas of the characteristic are overlooked, significantly limiting the meaningfulness of the test.
Criterion validity describes how well test results correspond to an external criterion. This validity is often assessed by calculating correlations between the test and the criterion. For example, the criterion validity of an employment aptitude test could be examined by comparing test results with the subsequent job performance of the people tested. A high correlation indicates good criterion validity, whereas a low correlation could cast doubt on the test’s validity.
Construct validity is particularly important because it checks whether a test actually measures what it claims to measure, rather than something else. It is determined using various methods, such as factor analyses or testing hypotheses. One example would be a self-confidence test: it should correlate positively with similar constructs such as self-esteem, but show little correlation with unrelated characteristics such as physical fitness. High construct validity increases confidence in the scientific value of a test and its suitability for diagnostic purposes.
