Learning objectives
By the end of this exercise, you should be able to:
- Correctly recode items (reverse coding).
- Calculate scale means and assess internal consistency using Cronbach’s alpha.
- Conduct an exploratory factor analysis (EFA) with three factors.
- Calculate descriptive statistics and interpret distributions.
- Identify skewed distributions and outliers.
Overview of the task
You have a dataset with 20 variables, including Likert scale data (measuring motivation, stress, and engagement) as well as some categorical variables (e.g., gender and field of study). The Likert items are assigned to three predefined scales, and some of them are reverse-coded.
Your task:
- Load the dataset (in R or SPSS)
- Explore the data structure (variable types, missing values, distributions)
- Perform reverse coding
- Calculate scale means
- Conduct an exploratory factor analysis (EFA) with three factors
- Calculate Cronbach’s alpha
- Identify skewed distributions and outliers
- Interpret the results
Data description
| Variable | Type | Description |
|---|---|---|
| motivation1, motivation2, motivation3, motivation4_rev | Likert (1-5) | Measures intrinsic motivation (motivation4_rev is reverse-coded) |
| stress1, stress2, stress3, stress4_rev | Likert (1-5) | Measures perceived stress (stress4_rev is reverse-coded) |
| engagement1, engagement2, engagement3, engagement4_rev | Likert (1-5) | Measures academic engagement (engagement4_rev is reverse-coded) |
| age | Numeric | Age of the participants |
| gender | Categorical | 1 = Male, 2 = Female, 3 = Diverse |
| study_program | Categorical | Study program |
| GPA | Numeric | Grade point average (0-4 scale) |
| social_media_hours | Numeric | Self-reported social media use per day (in hours) |
Step-by-step guide
- Load dataset
- Explore data
- Analyze distributions
- Analyze outliers
- Reverse coding (where necessary)
- Compute Cronbach’s alpha for all scales
- Calculate means
- Conduct an EFA
Solutions
First, we need to load the data.
read.csv()
Then we can take a closer look at the data. This is not strictly necessary, but it helps us get a feel for the data.
str(data)summary(data)colSums(is.na(data))
Of course, we are usually not interested in the raw data, but in aggregated data, which we should calculate. In the process, we may also need to recode some data.
data$motivation4_rev <- 6 - data$motivation4_revdata$stress4_rev <- 6 - data$stress4_revdata$engagement4_rev <- 6 - data$engagement4_rev
data$motivation_mean <- rowMeans(data[, c("motivation1", "motivation2", "motivation3", "motivation4_rev")])data$stress_mean <- rowMeans(data[, c("stress1", "stress2", "stress3", "stress4_rev")])data$engagement_mean <- rowMeans(data[, c("engagement1", "engagement2", "engagement3", "engagement4_rev")])
There are many methods available for the actual analysis, but here I’ll load the psych package..
library(psych)efa_result <- fa(data[, c(1:12)], nfactors = 3, rotate = "varimax")print(efa_result$loadings)
psych::alpha(data[, c("motivation1", "motivation2", "motivation3", "motivation4_rev")])psych::alpha(data[, c("stress1", "stress2", "stress3", "stress4_rev")])psych::alpha(data[, c("engagement1", "engagement2", "engagement3", "engagement4_rev")])
We can of course also approach the analysis visually.
library(moments)skewness(data$GPA)skewness(data$social_media_hours)boxplot(data$GPA, main="GPA distribution") boxplot(data$social_media_hours, main="Social media use")
Interpretation: factor loadings (which items load on which factors?), Cronbach’s alpha (is internal consistency adequate?); are there skewed variables or outliers?
- Open file
Analyze → Descriptive Statistics → FrequenciesCOMPUTE motivation4_rev = 6 - motivation4_rev. COMPUTE stress4_rev = 6 - stress4_rev. COMPUTE engagement4_rev = 6 - engagement4_rev. EXECUTE.COMPUTE motivation_mean = MEAN(motivation1, motivation2, motivation3, motivation4_rev). COMPUTE stress_mean = MEAN(stress1, stress2, stress3, stress4_rev). COMPUTE engagement_mean = MEAN(engagement1, engagement2, engagement3, engagement4_rev). EXECUTE.Analyze → Dimension Reduction → Factor AnalysisAnalyze → Scale → Reliability AnalysisAnalyze → Descriptive Statistics → Explore- Interpretation: factor loadings (which items load on which factors?), Cronbach’s alpha (is internal consistency adequate?); are there skewed variables or outliers?
Summary
This exercise combines practical data analysis with statistical concepts to foster a deep understanding of exploratory data analysis and factor analysis.
Alles klar?
Ich hoffe, der Beitrag war für dich soweit verständlich. Wenn du weitere Fragen hast, nutze bitte hier die Möglichkeit, eine Frage an mich zu stellen!
