Estimation: How We Determine Unknown Parameters in Statistics

In statistics, estimation is an indispensable tool for drawing conclusions about a population from samples. But how exactly does it work, and what are the most important concepts? In this blog post, I’ll explain the basics of estimation and give you practical examples that you can implement directly in R.

Point estimate: A specific estimate

Point estimation involves determining a single value that serves as the best estimator of an unknown population parameter. Let’s assume that you have a sample of $n$ observations $X_1, X_2, \dots$, and you want to estimate the population mean $\mu$. A well-known point estimator for this is the sample mean:

$\hat{\mu} = \frac{1}{n} \sum_{i=1}^{n} X_i$

Example: Point estimation in R

Suppose you have a sample of body heights (in cm), and you want to estimate the population mean. Here is an example of how you can do this in R:

# Sample of body heights
heights <- c(170, 165, 180, 175, 160, 185, 172, 168)

# Calculate the mean
mean(heights)

This code calculates the sample mean, which serves as an estimator of the population mean.

Properties of estimators

Not all estimators are equally good, which is why there are certain properties that a good estimator should satisfy:

  1. Unbiasedness: An estimator is unbiased if the mean of its sampling distribution corresponds to the true parameter. This means that, on average, the estimator produces the correct value.
  2. Consistency: An estimator is consistent if it approaches the true parameter value as the sample size increases.
  3. Mean Squared Error (MSE): The MSE indicates how far the estimates deviate from the true parameter value on average. It combines the bias and variance of the estimator.

Maximum Likelihood Estimation

One of the most common methods for estimating parameters is Maximum Likelihood Estimation (MLE). Here, you look for the parameter values that maximize the likelihood of the observed data. For many models, MLE leads to efficient and consistent estimators.

Example: MLE in R for a normal distribution

Suppose you have a sample drawn from a normally distributed population. You can easily calculate the maximum likelihood estimators for the mean and standard deviation of a normal distribution in R:

# Sample of normally distributed data
daten <- c(5.1, 5.5, 4.9, 5.0, 5.6, 5.2, 4.8, 5.4)

# Maximum likelihood estimation
logLik <- function(mu, sigma) {
n <- length(daten)
-n/2 * log(2*pi) - n/2 * log(sigma^2) - sum((daten - mu)^2) / (2 * sigma^2)
}

# Optimization to find the MLEs
mle <- optim(c(mean(daten), sd(daten)), function(par) -logLik(par[1], par[2]), hessian = TRUE)
mle$par # MLEs for the mean and standard deviation

In this example, you will find the maximum likelihood estimators for the mean and standard deviation of a normal distribution.

Interval Estimation: More Than Just a Point

The point estimate gives us a single value, but how certain are we about it? This is where confidence intervals come in. A confidence interval specifies a range that contains the true parameter with a certain probability (e.g., 95%). For the mean of a normally distributed population, the 95% confidence interval is:

$\hat{\mu} \pm z_{\alpha/2} \cdot \frac{\sigma}{\sqrt{n}}$​

Here, $z_{\alpha/2}$​ is the corresponding value of the standard normal distribution (e.g., 1.96 for 95%).

Example: Confidence Intervals in R

You can easily calculate a 95% confidence interval for the mean in R:

# Confidence interval for the mean
t.test(groessen)$conf.int

This code gives you the confidence interval for the mean of your sample.

Conclusion

Estimation is an essential part of statistics, and there are various approaches to estimating parameters. Point estimation provides a single estimate, while interval estimation gives you a range that is highly likely to contain the true parameter. Methods such as maximum likelihood estimation enable us to make precise estimates for more complex models. With R, you can implement all of these estimation methods quickly and easily.