Descriptive Statistics in R

This post is about descriptive statistics with R. You will learn:

  • What descriptive statistics actually is and why it is so valuable.
  • Which measures (e.g., mean, standard deviation) you can calculate in R.
  • Which R functions can help you get a good overview of your dataset.

(If you do not know much about statistics yet – no problem. This post will guide you through it step by step.)

A refresher: Why descriptive statistics?

Descriptive statistics is the art of describing and summarizing data – using measures (e.g., mean, median) and charts (e.g., boxplots).

Imagine this: You have 1,000 rows of data (a sample of 1,000 people). Who could still understand what the distribution looks like just by looking at it? Descriptive methods summarize everything so that you can quickly get a feel for where the data are concentrated and how much they vary.

These measures are not merely an end in themselves; they are essential for identifying errors, outliers, and peculiarities in the dataset and for developing hypotheses.

Measures of central tendency and dispersion in R

I assume you are already reasonably familiar with R. If not, take another look at the R basics.

Measures of central tendency

  1. Mean
    • Sensitive to outliers.
  2. Median
    • Divides sorted data into two equally sized halves.
    • More robust against outliers.
  3. Mode
    • The most frequent value.

R code examples:

werte <- c(5, 2, 9, 9, 7, 12, 2, 9)

mean(werte)    # Mittelwert

median(werte)  # Median

# Modus per Trick (die 'table'-Funktion)

modus <- names(sort(table(werte), decreasing=TRUE))[1]

modus

Measures of dispersion

  1. Standard deviation ($s$)
  2. Variance ($s^2$)
  3. Range (range = Max – Min)
  4. IQR (interquartile range = Q3 – Q1)

R code:

sd(werte)    # Standardabweichung

var(werte)   # Varianz

range(werte) # Spannweite

IQR(werte)   # Interquartilsabstand

Summary functions

  • summary(werte) gives you Min, 1st Qu., Median, Mean, 3rd Qu., and Max.

R code:

summary(werte)

#   Min. 1st Qu. Median  Mean 3rd Qu.  Max.

#    2       2     8     6.8     9     12

This gives you all the important numbers in one step.

Exploratory data analysis – practical workflow

A typical exploratory data analysis in R could look like this:

  1. Import your data, e.g. using read.csv(“myfile.csv”).
  2. Check the structure: str(data) and summary(data).
  3. Inspect missing values (NA), e.g. using is.na(data).
  4. Calculate descriptive statistics for each variable (mean, sd, etc.).
  5. Graphical checks – histograms, boxplots, etc. (more on this in the next blog posts).
  6. Investigate outliers.
  7. If necessary, apply a transformation (e.g. log(x)) if distributions are highly skewed.

Conclusion

You now have the basics of descriptive statistics in R at your fingertips:

  • You know the measures of central tendency and dispersion.
  • You know how to calculate them in R (mean(), median(), sd(), IQR(), etc.).
  • You can use summary() to quickly get an overview.
  • You have seen what an exploratory workflow can look like (import, summary, boxplot, and so on).

In this context, visualizations are also relevant, especially visualization with Base R: in other words, the classic R functions plot(), hist(), boxplot(), barplot(), and co.

Alles klar?

Ich hoffe, der Beitrag war für dich soweit verständlich. Wenn du weitere Fragen hast, nutze bitte hier die Möglichkeit, eine Frage an mich zu stellen!

Stelle Dominik eine Frage