Introduction to Statistics
In this post, we will focus a little more on statistics and quantitative research. The main focus is on applying statistics as part of a typical thesis at a university. We will begin with a general question: Why do we actually engage with statistics? Why should we learn about it? Second, we will broadly define what statistics is really about, and then move on to the more technical aspects.
First of all, we will talk about the paradigm. What kind of worldview do we assume when we apply statistics? Whenever we apply a method or theory, there must be something that goes beyond it, often referred to as a metatheory or paradigm, which determines the rules and tells us: “Is this really a good method for us to use, or might the method be being applied in a less-than-optimal way?”
We will also discuss the logic of knowledge generation that we apply in statistical research. This is primarily a deductive logic.
And finally, I’m using the word logic again here, but I mean something quite different. When we talk about statistics, there is usually something like a target population about which we want information, but at the same time, we often do not want to collect data from the entire population. Let me give you an example. In a newspaper, you might find a survey about the next election. Normally, this would not involve the election itself. That would be very expensive. The newspaper has only surveyed a few people. But how does this still allow us to draw valid conclusions about the target population? This is an important element when talking about statistics.
Why Learn Statistics?
Here, I’ll explain a few reasons why it is important to learn about statistics. I’d like to give you three reasons why I think statistics will be important for you and your life.
The first reason, of course, is that it can sometimes be very helpful to support your decisions with data, for example in business, but also in any other decision you make in life. There is some evidence behind your decisions, and statistics help you analyze that evidence in the right way. Sometimes, when you are dealing with numerical data, you need statistics and a general understanding of how statistics and probabilities work. That is the first reason why you should study statistics: it helps you make better, more informed decisions.
The second reason is that many people eventually write their own research paper. This could be a thesis , a report for your work, or anything else. But of course, you are required to go through the entire research process that we learn in this course in a systematic way that is transparent to others, that people can understand and identify with. So it is important not only to master the methods when applying them, but also to justify these methods and understand exactly what is happening, how to interpret the result, and so on.
Statistics are everywhere.
And finally, I would like to say that being familiar with statistics is actually a prerequisite for being a good citizen. Statistics are everywhere. Politics communicates a great deal with the help of statistics. Every news agency communicates extensively with the help of statistics. Statistical information surrounds us everywhere. And it is not at all easy to assume that everyone really understands this information. But if you do not understand this information, if, for example, you do not understand what politicians are telling you, what does that actually mean for you, for society, for your life? That is a problem, isn’t it?
If you are interested, I ask you to take a look at the case of Sally Clark . She was a lawyer in England who was charged primarily on the basis of statistical evidence, and in whose case laypeople and their understanding of statistics played a major role. And this trial really went in a very wrong direction; and perhaps a trigger warning here as well: It is a truly very depressing story. That is why I will not retell it here, but if you are interested, I can really recommend it; it is well documented online.
What is statistics, anyway?
I don’t think we have a problem talking about statistics, because statistics are truly everywhere, and I think that simply through coming into contact with them, we have developed a sense of what is and is not a statistic. But fundamentally, it is about a quantitative approach: a collection of numbers and other quantitative information that we first collect and then analyze numerically.
Put more formally, statistics is the science of the quantitative measurement and analysis of data. “Quantitative measurement” refers to quantifiable forms of expression for phenomena, such as the number of trainees in the IT sector. Mathematics serves as the language of mediation; the condition is that the phenomenon must also be expressible in numerical terms. “Quantitative analysis” refers to identifying trends, relationships, and so on with the help of statistical measures, guided by the research questions or hypotheses.
Statistics is the science of the quantitative measurement and analysis of data.
Langenscheidt Dictionary
Paradigm / Metatheory
Whenever we conduct research, we always have a set of assumptions about what the world we are studying actually is. We need these metatheories for our methods to work. These metatheories are also called paradigms.
When you engage in quantitatively oriented statistical research, you normally apply a paradigm called neopositivism. What does that mean? Basically, it means that there is such a thing as a reality out there that we can all perceive. It is a reality that we all look at. However, we do not always perceive this one reality in the same way.
It is a reality that we all look at. However, we do not always perceive this reality in the same way.
This results in a kind of probabilistic understanding of the world that we rely on with the help of statistics. You and I, we both look at the same object, the same reality, but we have our own perspectives, our own biases, and therefore there is only a certain probability that we will agree on specific characteristics of this reality.
Generate knowledge deductively
In science in general, there are various ways of generating knowledge through statistics. In statistics, we work with a very deductive logic, which means that we start with a theory, a kind of general knowledge, and then collect data and try to verify whether our theory still holds.
Overview of the research process
You can also find this deductive scheme in the following graphic about the research process:
Here, some of the steps are summarized once again—the four major steps that also structure the thesistribe program.
- Clarify the research model
- Clarify the research framework: From research topic to a more precise research question
- Clarify the theoretical foundations and derive hypotheses
- Collect data
- Operationalization, creating data collection instruments, quantification
- Conduct data collection
- Prepare the data: error checking and error correction
- Analyze data and test hypotheses
- Creating indices, item analyses, scale scores; Univariate statistics
- Difference and association analyses
- Interpreting results
- Reporting
- Introduction
- Background
- Method
- Results
- Discussion
Exploration vs. confirmation
We therefore have an approach that is also referred to as a confirmatory method. And that is very important, because it means you cannot address research questions that are highly exploratory.
If we do not have a theory at the outset, the methods we have learned about in the following sections cannot be applied. We need a theory.
So we have a suitable research question, and we break it down to make the whole thing manageable. And that means two things. First, we have to put it in a box. We call this a statistical model. We are not interested in reality in all its richness. We are only interested in a tiny area, which we call a statistical model and in which our variables are located.
For example, if you are interested in how gender affects income, then gender is relevant here. There is also another concept, namely income. For example, your level of competence or something similar, which would of course affect income level.If it is not explicitly included in our model, it does not exist for us or for the statistical methods we apply. So it is very important to have a good model.
And then we also have to think about how we can measure this. There is a step called operationalization, which basically means that you have to make things measurable. And that is a very, very difficult thing, right? Let me tell you that straight to your face: It is difficult to measure some concepts. Others are easier.
For example, if you say that you are interested in gender as defined in someone’s passport, that is a relatively easy variable to study, because you can simply look at people’s passports and see it with your own eyes, right? But what if, for example, you have a more nuanced understanding of gender, or if it is about sexual orientation, personality, or something similar? Then it is not directly observable. That is what is known as a latent variable.
That is why we all need to discuss at length how such things can be measured, because it is definitely not that simple. And when we are reasonably confident that what we have done is more or less acceptable—that it constitutes a valid, reliable, and objective measurement—then we can actually collect the data and analyze it using statistical samples.
Here, we will get to know many different tests that we can use. They are all based on your theory, your research question, and the hypotheses that you derive from your research question. And when we have a result—usually just one or two numbers—we try to interpret those numbers, write a report, and that’s it. That is the logic.
You can see that there is a very deductive logic at work: first, we model a great deal of theory, then we collect the data, and finally we compare the two. Just to create a small contrast: if you conduct a highly exploratory interview study—for example, a survey of experts in the field—you may sometimes have no theory at all. You are simply interested in a particular area. You go out, talk to people, and may not even know what the concept you want to investigate means. So you discuss it with others to find out. At the very end of such a research project, you may develop a hypothesis. This is often referred to as a hypothesis-generating approach. Our statistics course, on the other hand, is more about testing and confirming hypotheses.
Whatever we do in the following sections, whatever project you carry out with the help of statistics, try to remain aware of each of these steps and whether they are actually present. There needs to be this kind of linear direction for our model of knowledge generation to work properly.
The logic of inferential statistics
What is statistics fundamentally about? Let’s think of an example. Let’s think about an election and a prediction about an election.
It is quite costly to ask everyone—as is done in the election itself. That is why we do not normally hold polls every day. What we do encounter every day, however, are statistics that make predictions about elections, and these are often quite accurate as well. So how do they work?
So, let’s assume that there is a group of people who can vote, known as the population or target population. From this population, we try to draw what is known as a sample. During sampling, people are selected from the population and included in the specific sample that we want to survey. We then ask this sample a question, for example: “If you were voting today, who would you vote for?”
Once we have this data, we can draw an interesting inference about the population, because normally we are not satisfied with merely describing what the result would look like if this particular sample were to vote. We are much more interested in making what is known as an inference about the population and saying: Okay, if there were a vote today, we would expect the entire population to vote in this way. In this way, the focus is no longer so much on a particular sample, but on the population as a whole. And that is the great thing about statistics: We can draw conclusions about the entire population at a fraction of the cost, because we only pay for the smaller sample we surveyed. And those costs can be remarkably low.
So it is really cost-effective to use statistics. But that also means that I always think of it as a kind of closed loop. On the one hand, you have the population and draw a sample from it, and then you make an inference to move from the sample back to the population. This means that the sampling procedure has to be quite robust.
Summary
Purpose of Statistics
What is statistics about? I believe that, ultimately, statistics is about decision-making—and we make decisions quite often. We probably make a few thousand decisions every day. Some of them may not be very important, but when it comes to the more important ones, we want to make sure that we can make a good decision and that there is perhaps some evidence to support it, so that we feel more confident. This is where statistics plays a very important role.
To give you a slightly different perspective: it is about seeing something that our eyes are not very good at perceiving. If you imagine having a large pool of data in front of you, it is often very difficult to recognize what the numbers are actually trying to tell you. But statistics makes the picture clearer. Statistics gives you a pair of glasses that allows you to see the patterns underlying your dataset, and this will ultimately help you make better decisions.
Meta-theories of statistics
As for the metatheory, or paradigm: What rules apply to statistics? This is often referred to as the neo-positivist paradigm. You may remember that it is based on the idea of an objectively real world. Reality exists and is basically the same for everyone, but with one caveat. That is, we cannot perceive it in the same way because there is a subjective level, subjective bias, and so on. As a result, we can see only an incomplete and distorted version of this objective reality. At the very least, this is something to keep in mind when interpreting the data.
Maybe you have already worked with other methods—perhaps you have conducted reconstructive research, for example through narrative interviews in which people talk extensively about their careers, their lives, and everything else. Your approach to this question would be quite different, wouldn’t it? This is very important, especially if you already have experience with other methods, because you may then need to make a somewhat greater mental distinction between them.
The Knowledge-Generation Process
And then we also talked about the process of creating knowledge. In statistical research, deductive logic means that we first have a theory or a general understanding of how something works, and then try to apply it to a specific observation. This specific observation is essentially your empirical dataset. To remind you: deductive logic is fairly straightforward.
For example, if you look at a quantitative research paper, you will see that it includes something like a theoretical background from which the hypotheses are derived, right? The theory is explained. That is, so to speak, the starting point. Then there is a methods section describing how the data are collected and analyzed, and finally the results section, in which the empirical data are presented. So this is exactly the logic I have just outlined. It is deductive logic: first, we have general knowledge—the theory—and then we have the empirical results, the more specific knowledge, you might say. At the end, the question is: How well do they fit together? Is what we knew beforehand—the theoretical part—more or less consistent with what our empirical data show? This is what is called the discussion in a research paper.
Outlook: Inferential Statistics
And finally, we discussed conclusions or inferences. We will talk about that in more detail in a later section. But it was important to me to sketch this out here at the beginning, because I believe it is quite important for understanding what we are actually doing here.
As a reminder: I have always imagined it as a circle. We have this target population that we are really interested in. But because we may not be able to collect that much data, or perhaps cannot reach the people at all, we ask only a few of them—or observe a few of them. Whatever your data collection methods look like, we will talk about that shortly as well. We use a sample, as it is called. This is a sampling method that we apply here in order to make the data as unbiased and balanced as possible. Representative is another word we often hear, and once we have some results at the sample level, from a small subsample of the overall population, we still try to make an estimate and draw conclusions about the overall population.
And that is exactly what is so beautiful about statistics. I believe it gives the methods we work with a great deal of power. At the same time, of course, it reminds you to be very careful when drawing your sample. You need a very good sample; otherwise, your data could actually be misleading.
