Sample and population in statistics
When we talk about statistics, we often come across the terms “sample” and “population.” These two terms are essential for understanding data analysis, especially in psychology and occupational psychology. Let’s clearly distinguish between them and also discuss the element that connects them: the sampling procedure.
What is a population?
A population, also known as the total population, is the complete set of all possible objects about which you want to make a statement. This could be, for example, a group of people, companies, behaviors, or even psychological characteristics. When we consider a population in psychology, this could mean, for instance: “all students at a university,” “all employees in a company,” or “all people who work under certain working conditions.”
It is important to note that a population is often very large and, in practice, it is frequently impossible to study it in its entirety. Imagine you want to find out how stressful work is in Germany. The population would then consist of everyone who works in Germany. It is obvious that surveying all of these people would hardly be feasible—which is where the sample comes in.
What is a sample?
A sample is a subset of the population that you actually study. Ideally, it should be selected so that it accurately represents the population. The quality of the sample directly affects the meaningfulness of your analysis. If the sample is representative, you can use the results of your study to draw conclusions about the population as a whole.
For example, imagine you want to know how satisfied employees at a company are. It may not be practical to survey every employee, so you select a random sample of 100 employees. If this sample is well selected, you can assume that the satisfaction levels of these 100 employees provide a reliable indication of the satisfaction of all employees at the company.
Why do we need samples?
A complete study of the population (known as a census) is in most cases too time-consuming, too expensive, or simply impossible. By working with samples, we can still make well-founded statements about the population. It is important, however, that the sample is selected randomly and representatively in order to avoid bias.
Here is an example from psychological research: Suppose you want to compare the stress levels of employees in different industries. The population would consist of all employees in each industry. Since it is practically impossible to survey every employee, you select samples from each industry and then compare the stress levels across these samples.
Distinguishing Between Data and Population Parameters
We now turn to the distinction between data and population parameters. We are often interested in describing specific characteristics of a population, such as the mean or variance of a variable. However, since we usually only study samples, we have to estimate these values based on the sample.
Population Parameters
A population parameter is a measure that describes the entire population. Typical population parameters are the population mean μ or the population variance σ². These values are generally unknown because we cannot study the entire population.
Sample Statistics
A sample statistic is a measure calculated from a sample. For example, the sample mean x̄ is an estimate of the population mean μ (μ is pronounced “mu”). Similarly, the sample variance s² is an estimate of the population variance σ² (σ is pronounced “sigma”). We use sample statistics to draw conclusions about population parameters.
Example of the Difference Between a Sample Statistic and a Population Parameter
Suppose you are studying average job satisfaction in a company. The actual job satisfaction level of all employees would be the population mean μ. However, since it is not possible to survey all employees, you survey a sample and calculate the mean x̄ for that sample. This sample mean is an estimate of the population mean.
When we use sample statistics, we need to be aware that there is always an element of uncertainty. We can quantify this uncertainty using statistical methods such as confidence intervals or hypothesis tests.
Sampling Methods
The way in which a sample is drawn from a population is crucial for the validity of the conclusions we can draw from the data. There are various sampling methods, which should be selected depending on the research objective, available resources, and specific circumstances.
Simple Random Sampling
The most common method is simple random sampling. Here, every member of the population has an equal probability of being selected. This method minimizes bias and increases the likelihood that the sample is representative of the population. For example, if you have a list of all employees in a company, you could use a random mechanism (such as a computer program) to decide who is included in the sample.
Stratified Sampling
In stratified sampling, the population is divided into subgroups, known as strata, that differ in certain characteristics (e.g., age, gender, or department). A random sample is then drawn from each stratum. This ensures that every relevant subgroup is adequately represented in the sample. For example, you might want to ensure that both male and female employees are represented in your sample in proportion to their share of the total population.
Cluster Sampling
In cluster sampling, the population is divided into groups, or “clusters,” and some of these clusters are then selected at random. All members of a selected cluster are included in the sample. This method is often used when it is logistically difficult to conduct simple random sampling—for example, if you want to study schools and survey a random selection of schools instead of individual students.
Quota Sampling
Quota sampling is similar to stratified sampling, but no random selection is made from the strata. Instead, a fixed number of people (a “quota”) is specified for each stratum, and you actively select participants until this quota is reached. This can lead to an unrepresentative sample because there is no element of randomness. Nevertheless, this method makes it possible to deliberately include specific groups in a predetermined number. For example, you might want to ensure that your sample contains 50 men and 50 women by actively searching until you have met these quotas.
Convenience Sample (Convenience Sampling)
Convenience sampling is a method in which the sample is selected based on the easy availability and accessibility of people. This means that you select the people who are easiest to reach, without regard for representativeness or randomness. This method is quick and inexpensive, but it carries a high risk of bias because the results may not be generalizable to the entire population. A typical example would be surveying your colleague near your office instead of randomly selecting employees from across the entire organization.
Self-assessment questions
The population comprises all units about which a statement is to be made, e.g., all residents of a country.
A sample is a subset of the population that is studied in order to draw conclusions about the population as a whole.
Why use samples?
It is often impractical or impossible to study the entire population because this would be too costly, time-consuming, or inaccessible. Samples make it possible to draw reliable conclusions about the population with limited effort, provided that they are representative.
Example: In a survey on voting preferences, a sample of 1,000 voters is selected to predict the preferences of the population as a whole.
Representativeness means that the sample reflects the characteristics of the population as accurately as possible.
This can be ensured through:
- Random sampling: Every unit has an equal chance of being included in the sample.
- Stratification: The population is divided into strata (e.g., by age or gender), and a sample is drawn from each stratum in proportion to its size.
- Adequate sample size: A sample that is too small increases the risk of bias.
Example: In a university survey, stratifying by field of study could ensure that all academic disciplines are represented.
Random sample:
- Advantage: Objective and representative selection.
- Disadvantage: Can be time-consuming if a complete list of the population is unavailable.
Stratified sample:
- Advantage: Ensures that relevant subgroups are represented.
- Disadvantage: Requires precise knowledge of the population.
Cluster sample:
- Advantage: Cost-effective, as groups (e.g., classes or cities) are studied as a whole.
- Disadvantage: Higher risk of bias if the clusters differ substantially.
Convenience sample:
- Advantage: Simple and inexpensive.
- Disadvantage: Usually not representative.
