Exploratory Factor Analysis in R

Exploratory factor analysis (EFA) is one of the most important techniques in psychology and the social sciences for identifying latent constructs in data. It helps us understand the structure of questionnaires or tests and eliminate irrelevant items. In this blog post, I’ll show you how to conduct an EFA in R, including data preparation, factor extraction, rotation, and interpretation of the results. In addition, we’ll explore the methodological background and present further examples of applications in various research fields.

📌 Definition: Exploratory factor analysis is a statistical method for extracting underlying latent variables (factors) from a large number of observed variables. It is primarily used to identify structure and differs from confirmatory factor analysis (CFA), which tests hypotheses.

Data Preparation for EFA

Before conducting a factor analysis, we need to make sure that the data meet the necessary requirements:

Handling Missing Values

Missing values can distort the analysis and lead to incorrect conclusions. A simple method is mean imputation, which can, however, introduce bias. A more advanced technique is multiple imputation, in which missing values are simulated multiple times and the results are averaged. This method is based on probability models and is particularly useful when many values are missing. However, it requires careful model selection and is a more advanced method that goes beyond simple imputation strategies.

library(mice)
data <- mice(data, method = 'pmm', m = 5)
data <- complete(data)

Bartlett’s test of sphericity

Tests whether the correlation matrix corresponds to an identity matrix. If so, the variables are uncorrelated, which argues against conducting a factor analysis. A significant Bartlett’s test indicates that the correlations in the matrix are sufficiently strong to make factor analysis meaningful.

library(psych)
cortest.bartlett(cor(data), n = nrow(data))

Kaiser-Meyer-Olkin (KMO) test

Assesses the suitability of the data for factor analysis by measuring how well each variable is represented by the extracted factors. The KMO value ranges from 0 to 1. A value above 0.8 is considered excellent, 0.7 good, and below 0.5 unsuitable for factor analysis. The test is based on calculating partial correlations: if the partial correlations are small, this indicates that factor analysis is appropriate because the shared variances between the variables are high. The KMO test is an advanced method for assessing data quality.

KMO(r = cor(data))

Check the level of measurement

EFA assumes that the variables are measured at the interval level.

Avoid multicollinearity

A high correlation between variables can lead to bias. Calculating the determinant of the correlation matrix helps identify multicollinearity :

det(cor(data))

Factor Extraction

The most common methods of factor extraction are principal component analysis, principal axis factoring, and maximum likelihood factor analysis.

Principal Component Analysis (PCA)

PCA is a method for reducing the dimensionality of data by transforming the original variables into a new set of uncorrelated variables, known as principal components. These principal components are linear combinations of the original variables and capture the maximum amount of the data’s total variance. The first principal component contains the most variance, the second contains the second-most, and so on. This method is frequently used to simplify datasets with many variables and make them more manageable visually and computationally without losing essential information.

Principal Axis Factoring (PAF)

PAF is a factor extraction method used specifically to identify latent constructs. It is based on the assumption that observed variables are influenced by a combination of latent factors and random measurement errors. Unlike principal component analysis, which considers the total variance of the variables, PAF analyzes only the common variance (shared variance) of the variables. Iterative procedures are used to estimate factor loadings and separate error variance from total variance. This makes PAF particularly suitable for psychological and social science research questions in which latent constructs play a central role.

Maximum Likelihood Factor Analysis

This procedure is based on maximum likelihood estimation to optimize the factor solution. It enables not only the estimation of factor loadings, but also statistical tests of whether the selected number of factors is appropriate. These include tests of model fit, such as the likelihood-ratio test, which examines whether a factor solution with fewer factors is significantly worse than a more comprehensive solution. The method assumes that the data are multivariately normally distributed, which in practice is often only approximately true. It is particularly useful in inferential statistical contexts where formal hypothesis testing is desired, but compared with simpler methods such as principal axis factoring (PAF), it requires a more careful examination of the underlying assumptions.

Here is a PAF using the psych package:

fa_model <- fa(data, nfactors = 3, fm = "pa", rotate = "none")
print(fa_model)

📌 Definition: Factor extraction is the mathematical process of identifying latent variables from patterns of correlations among observed variables.

An alternative method for determining the number of factors is parallel analysis, This method compares the eigenvalues (a measure of the variance explained by a factor in the data) of the empirical data with those of randomly generated data sets of the same size. Factors are retained when their eigenvalues are greater than the randomly generated values. Parallel analysis thereby avoids overestimating or underestimating the number of factors, which is particularly important in psychometric studies. While the Kaiser criterion (eigenvalue > 1) and the scree test are often used, parallel analysis provides a more robust, statistically grounded basis for decision-making.

fa.parallel(data, fa = "fa")

Factor rotation

Because unrotated factors are often difficult to interpret, we apply a rotation. The two main types are orthogonal and oblique rotation.

📌 Definition: Factor rotation improves the interpretability of factors by transforming the loadings so that each variable correlates strongly with one factor.

Orthogonal Rotation (Varimax)

This factor rotation method ensures that the extracted factors remain uncorrelated, meaning that they are independent of one another. The aim is to transform the factor loadings so that each variable loads as strongly as possible on a single factor and has as low a loading as possible on the other factors. This improves interpretability because each variable can be clearly assigned to a specific factor. Varimax rotation is one of the most frequently used orthogonal rotation methods and is based on maximizing the variance of the squared loadings within each factor. This means that high loadings become even higher and low loadings become even lower, enabling a clearer structure of the data. However, this method has the limitation that it assumes the underlying factors are actually uncorrelated, which is not always the case in practice.

Oblique Rotation (Oblimin)

This method allows factors to influence one another, which is more plausible for many real-world psychological and social science datasets than assuming complete independence (as with Varimax). Oblimin maximizes the simple structure of the factor loadings and allows for better interpretation of the factors when they are not completely uncorrelated. The degree of correlation between the factors can be controlled using the so-called delta parameter. A low delta value (e.g., 0) leads to greater independence between the factors, whereas a higher delta value allows for stronger correlation. This method is particularly helpful when one assumes that the underlying latent variables are conceptually related, as is often the case with psychological constructs.

Here is an example of an oblique rotation:

fa_model_rot <- fa(data, nfactors = 3, fm = "pa", rotate = "oblimin")
print(fa_model_rot)

Interpretation of the Factors

After the rotation, we interpret the factors based on the factor loadings:

  • Loadings > |0.5|: Strong association with a factor.
  • Loadings between |0.3| and |0.5|: Moderate association.
  • Loadings < |0.3|: Unimportant items.

A scree plot helps determine how many factors should be retained:

fa.parallel(data, fa = "fa")

A biplot can also be helpful:

fa.diagram(fa_model_rot)

Examples of Applications

Example from Psychology

A questionnaire on job satisfaction is analyzed using EFA. Two factors emerge:

  • Factor 1: Satisfaction with the Work
  • Factor 2: Satisfaction with Pay

Example from Business

A customer survey on product satisfaction reveals the following factors:

  • Factor 1: Product Quality
  • Factor 2: Customer Service

Conclusion

Exploratory factor analysis is a powerful tool for exploring latent structures in data. With R, you can implement it efficiently and interpret the results visually.

Key points:

  1. Prepare and check the data.
  2. Choose a suitable extraction method.
  3. Use rotation for better interpretation.
  4. Use factor loadings and visualizations.