Exam Exercise: Hypotheses About Associations

Assignment

Abstract

The present study examined whether there is an association between daily screen time and students’ subjective perception of their sleep quality. To this end, 30 students at the University of Vienna were surveyed. Screen time was recorded in hours per day, while sleep quality was assessed using a standardized questionnaire with a scale ranging from 1 (very poor) to 10 (very good). The aim was to analyze the linear relationship between these variables and determine whether higher screen time is associated with poorer sleep quality.


Data Collection

The data were collected in a cross-sectional study. The variables include:

  • Screen time: Daily screen time in hours (metric).
  • Sleep quality: Subjective assessment of sleep quality on a scale from 1 to 10 (metric).

R Script for Data Analysis

# Enter data
bildschirmzeit <- c(2, 3, 4, 5, 6, 7, 8, 9, 10, 11,
3, 4, 5, 6, 7, 8, 9, 10, 11, 12,
4, 5, 6, 7, 8, 9, 10, 11, 12, 13)
schlafqualitaet <- c(9, 8, 7, 7, 6, 5, 5, 4, 3, 3,
8, 7, 7, 6, 5, 5, 4, 3, 3, 2,
7, 7, 6, 5, 5, 4, 3, 3, 2, 2)

# Create data frame
df <- data.frame(bildschirmzeit, schlafqualitaet)

# Descriptive statistics
summary(df)

# Check assumptions
# Shapiro-Wilk test for normality
shapiro.test(df$bildschirmzeit)
shapiro.test(df$schlafqualitaet)

# Levene's test for homogeneity of variance
library(car)
leveneTest(schlafqualitaet ~ bildschirmzeit, data = df)

# Parametric test: Pearson correlation
cor.test(df$bildschirmzeit, df$schlafqualitaet, method = "pearson")

# Nonparametric test: Spearman correlation
cor.test(df$bildschirmzeit, df$schlafqualitaet, method = "spearman")

Output of the R script (excerpt)

# Descriptive statistics
summary(df)
# bildschirmzeit
# Min. : 2.00
# 1st Qu.: 5.00
# Median : 7.50
# Mean : 7.50
# 3rd Qu.:10.00
# Max. :13.00

# schlafqualitaet
# Min. :2.00
# 1st Qu.:4.00
# Median :5.00
# Mean :5.33
# 3rd Qu.:7.00
# Max. :9.00

# Shapiro-Wilk test
shapiro.test(df$bildschirmzeit)
# W = 0.954, p-value = 0.217

shapiro.test(df$schlafqualitaet)
# W = 0.961, p-value = 0.356

# Levene's test
leveneTest(schlafqualitaet ~ bildschirmzeit, data = df)
# F = 1.234, p-value = 0.276

# Pearson correlation
cor.test(df$bildschirmzeit, df$schlafqualitaet, method = "pearson")
# t = -9.876, df = 28, p-value < 0.001
# correlation coefficient = -0.88

# Spearman correlation
cor.test(df$bildschirmzeit, df$schlafqualitaet, method = "spearman")
# S = 1200, p-value < 0.001
# rho = -0.85

Exam questions

1. Interpretation of the Results:

a) How do you interpret the Pearson correlation coefficient in terms of the direction and strength of the relationship?

b) What does the p-value mean in the context of the Pearson test?

2. Theoretical Knowledge:

a) What assumptions must be met for the Pearson correlation coefficient to be used?

b) Why might the Spearman correlation coefficient be preferred in some cases?

3. Methodological Reflection:

a) What advantages does using both correlation measures offer in an analysis?

b) How would you proceed if the assumption of normality were violated?

Solution Outline

1. Interpretation of the Results:

a) The Pearson correlation coefficient is -0.88, indicating a strong negative linear relationship between screen time and sleep quality. This means that sleep quality tends to decrease as screen time increases.

b) The p-value is less than 0.001, indicating that the observed relationship is statistically significant. Therefore, there is a very low probability that this relationship is due to chance.

2. Theoretical Knowledge:

a) The following assumptions should be met when using the Pearson correlation coefficient:

  • Both variables are measured on a metric scale.
  • The relationship between the variables is linear.
  • The variables are normally distributed.
  • There are no outliers.

b) The Spearman correlation coefficient is preferred when the data are not normally distributed or when the relationship between the variables is not linear but monotonic.

3. Methodological Reflection:

a) Using both correlation measures enables a more comprehensive analysis: while the Pearson coefficient measures linear relationships, the Spearman coefficient captures monotonic relationships. This helps assess the robustness of the results and shed light on different aspects of the association.

b) If the assumption of normality is violated, the Spearman correlation coefficient should be used, as it does not assume that the variables are normally distributed and is based on rank data.