Why Bayesian Statistics?
Bayesian statistics is an alternative approach to statistical inference that differs from traditional frequentist statistics. It allows uncertainties to be modeled explicitly and prior knowledge to be incorporated into the analysis.
Definition: Bayesian statistics is based on probability theory and uses Bayes’ theorem to update knowledge with new data.
Differences Between Frequentist and Bayesian Statistics
| Property | Frequentist Statistics | Bayesian Statistics |
|---|---|---|
| View of Parameters | Fixed but unknown value | Probability distribution |
| Inference | Based on hypothetical samples | Combination of prior knowledge & data |
| Confidence Intervals | Statement about hypothetical repetition | Direct probability |
Example: Medical Test
Suppose a test for a rare disease is 99% accurate. A patient tests positive. How likely is it that they actually have the disease? Bayesian statistics provides a systematic way to solve this problem.
Bayes’ Theorem
The central element of Bayesian statistics is the Bayes theorem:
$ P(\theta | D) = \frac{P(D | \theta) P(\theta)}{P(D)} $
Meaning of the terms:
- Prior $P(\theta)$: Prior knowledge about the parameter $\theta$
- Likelihood $P(D | \theta)$: Probability of the data given $\theta$
- Posterior $P(\theta | D)$: Updated distribution of the parameter after taking the data into account
- Normalizing factor $P(D)$: Overall probability of the data
Visual example: A biased coin?
Imagine you have a coin that you suspect is not fair. Before tossing it, you believe that it is probably fair, but not with certainty. You toss it 10 times and observe heads 8 times.
Step 1: Choosing the prior
Suppose we model our uncertainty with a Beta distribution: $ P(\theta) \sim \text{Beta}(2,2) $
Step 2: Calculating the likelihood
The probability of getting heads 8 out of 10 times depends on the true $\theta$ value.
Step 3: Calculating the posterior distribution
The new distribution is obtained by multiplying the prior and the likelihood: $ P(\theta | D) \propto P(D | \theta) P(\theta) $
After applying Bayes’ theorem, we obtain an updated Beta distribution:
$ P(\theta | D) \sim \text{Beta}(10,4) $
Conjugate priors
Definition: A prior is conjugate if the posterior distribution has the same form as the prior distribution.
| Likelihood | Conjugate Prior | Posterior |
|---|---|---|
| Binomial | Beta | Beta |
| Normal (mean known) | Normal | Normal |
| Poisson | Gamma | Gamma |
Example: Binomial Distribution with Beta Prior
Assume that a company wants to estimate the success rate of its new advertising campaign. It starts with a Beta(2,2) prior, assuming that success is more likely than 50%, but not extremely likely.
After 100 campaigns, 70 are successful: $ P(\theta | D) \sim \text{Beta}(2+70, 2+30) $
Credible Intervals
Definition: A credible interval indicates the probability that the true parameter value lies within a specific range.
Unlike the frequentist confidence interval, which refers to hypothetical samples, a 95% credible interval directly indicates that the true value lies within this range with a 95% probability.
R Code for Illustration
Here is some R code to visualize a beta distribution before and after updating with data:
# Installiere und lade das ggplot2 Paket
install.packages("ggplot2")
library(ggplot2)
# Prior- und Posterior-Verteilungen erzeugen
x <- seq(0, 1, length=100)
prior <- dbeta(x, 2, 2)
posterior <- dbeta(x, 10, 4)
data <- data.frame(x, prior, posterior)
# Plot
ggplot(data, aes(x)) +
geom_line(aes(y = prior), color = "blue", linetype = "dashed", size = 1.2) +
geom_line(aes(y = posterior), color = "red", size = 1.2) +
labs(title = "Bayessche Aktualisierung: Prior vs. Posterior",
x = "Wahrscheinlichkeit für Kopf",
y = "Dichte")
Alles klar?
Ich hoffe, der Beitrag war für dich soweit verständlich. Wenn du weitere Fragen hast, nutze bitte hier die Möglichkeit, eine Frage an mich zu stellen!
