Measurement and Levels of Measurement in Statistics

In the world of research, particularly in statistics, measurement plays a fundamental role. It enables us to quantify the properties or characteristics of objects and phenomena and thus make them comparable. However, not all measurements are the same. The way data are measured and interpreted depends on the measurement level or scale level. In this blog post, we take a closer look at the different measurement levels and their significance in research.

Fundamentals of Measurement in Statistics

Measurement means assigning numbers to objects or events according to established rules. This assignment should preserve the structure of the objects or events; in other words, the relationships between the numbers should reflect the relationships between the objects being measured. A simple example is assigning numbers to genders (e.g., female = 1, male = 2). However, these numbers convey different information depending on the underlying measurement level.

statistics

Levels of Measurement

Measurement levels can be divided into four main categories: nominal, ordinal, interval, and ratio scales. Each of these levels offers different possibilities for interpreting data.

Nominal Scale

The nominal scale is about categories. Data at this level can be divided into different groups, with each group represented by a unique number. The only meaningful relationship between these numbers is equality or inequality. Examples include religion, citizenship, and gender. The nominal scale is characterized by disjointness and exhaustiveness: each characteristic is assigned a unique code, and all possible values of the characteristic are captured.

Ordinal scale

The ordinal scale extends the concept of the nominal scale by introducing a ranking among the categories. Here, it is not only category membership that matters, but also the relative position within a series of categories. Examples include Likert scales and school grades. Although we know that a grade of “very good” is better than “good,” we cannot quantify how much better it is.

Interval scale

With the interval scale, the differences between measured values are meaningful. This means that we know not only the order of the values, but also the exact distance between them. A classic example is temperature measured in degrees Celsius. The interval scale allows us to perform operations such as addition and subtraction meaningfully, but it does not have a natural zero point.

Ratio scale

The ratio scale provides the most information and is the most advanced scale. In addition to the properties of the interval scale, the ratio scale has an absolute zero point, which indicates the absence of the measured characteristic. This makes it possible to establish ratios between measured values. Examples include income and body height. All mathematical operations are meaningful on this scale.

The significance of levels of measurement in research

Understanding the different levels of measurement is crucial for the correct analysis and interpretation of data in research. Each level offers its own possibilities and limitations, which must be taken into account when collecting, analyzing, and interpreting data. By choosing the appropriate level of measurement, researchers can ensure that their studies produce valid and meaningful results.

Carefully selecting and applying the different levels of measurement enables us to better understand and explain the complex world around us. Whether the goal is to classify social phenomena, measure attitudes, or quantify physical properties, knowing and applying the right levels of measurement is an indispensable tool in research.

Once again, in the original wording

Understanding levels of measurement is very important. And sometimes it is easier to follow the whole thing when it is spoken. So here is a transcript of one of my lectures on statistics:

Now I am focusing on the levels of measurement. And this is a pretty important topic. It is an important topic because it gives you an idea of the quality of the data. And what is even more important is that it will be very, very important when we deal with statistical tests. Because I can tell you one thing: You will find it easy to carry out any statistical test. You will be able to perform these tests, and I will also give you enough guidance to interpret them. But one challenge I sometimes see among students is that they do not know which test they should actually choose. That is because we will get to know quite a number of them. There are between ten and 20 tests that we will cover at the core of this course. Knowing about levels of measurement will help you make this decision.

In fact, for the area of statistics we are discussing here, you will get a definitive answer as to which test you should choose once you know the levels of measurement. There are four that you need to know.

The first is the nominal scale, which we will come to in a moment. Then there is the ordinal scale, and finally the interval and ratio scales. In most statistical programs, such as SPSS, the last two are simply combined. They call this the level of measurement or the metric level of measurement, and they mean both of them. For the purposes of applying statistics, it is therefore perfectly fine to focus on these three levels: nominal, ordinal, and metric.

Let us start with nominal data. The name says it all. This is essentially all we have: categories based on names. We can say that someone belongs to a particular category or does not belong to that category. For example, when it comes to gender, there are the categories male, female, and diverse. These are labels that you can use, and you can assign people to these categories, but you cannot do anything else with them. Of course, because we work with computers, we do use numbers for these categories. 1 is male, 2 is female, and 3 is diverse. Nevertheless, we cannot perform calculations with these numbers. There is no order in the data. No calculation would be a valid operation here, because you are not dealing with actual numbers; these numbers are merely symbols representing text.

That was the first level of measurement scales: the nominal level of measurement. If you want to go one level higher, you have ordinal data. And with ordinal data, the function is already apparent in the name—you can order the data. This suddenly gives you a kind of structure. In psychological research, Likert scales are often used, for example—scales that indicate how strongly you agree with a particular statement. Not at all, A little, Very much. Based on these data, you can say that there is a certain order.

But you still cannot do an awful lot with this order; for example, you cannot simply perform addition or subtraction. This is because the distances between these categories do not have to be equal. Here is an example, although it depends somewhat on the education system in which you grew up: Many education systems use a grading system with ordinal values. You could have grades A, B, C, D, E, and F, for example. Of course, we could also simply use the numbers 1–6 for them.

F could be the failing grade, while the other five could be the passing grades, for which you actually pass the test. But F usually covers 50% of the entire range, right? So it does not represent the same interval. F could mean that you achieved 0% on the test, or it could mean that you achieved 50%, and the range becomes much smaller as you move up through the letters.

Okay, so here too, you cannot really perform calculations. You can only arrange the values in an order. Yes, D is a better grade than E and F. B is a better grade than D. But you cannot say that the difference between A and B is the same as the difference between E and F, for example.

So if you want to move up one level, we come to data on the interval scale. The classic example here is temperature, particularly the Celsius or Fahrenheit scale, because the special feature of this scale is that you can perform certain calculations. You can add temperatures, and you can subtract temperatures. You could say that four degrees Celsius plus two degrees Celsius equals six degrees Celsius.

And we can do this because the intervals between one degree Celsius or one degree Fahrenheit are, of course, always the same size. So I can perform this calculation. What I cannot do, however, is multiply or divide, because there is no natural baseline; there is no natural zero point. So 40 degrees Celsius is not twice as warm—or would that be cold? I don’t know—as 20 degrees. This calculation doesn’t work because you need a natural zero point for that. Think about your income. You can have an income of zero, an income of 1,000 units, or an income of 2,000 units. And if you have that, you can say, yes, 2,000 is exactly twice as much as 1,000. So you can establish these ratios. Here, too, the name is very meaningful: this is a ratio scale.

Those were the four levels that you should definitely remember. And I want you to see the hierarchy in which they are arranged. If you think of data quality, you could say that the nominal level is a fairly low level of quality, while the ratio level represents, so to speak, the highest level of quality. And if you have data at a higher level, you can always convert it to a lower level. To give you a simple example: let’s say I have a survey in which I ask about your age.

So please tell me your age in years. Okay, you give me a number and say it is 25. So this number is definitely a metric value: 25. There is a natural zero. This means that 50 is exactly twice as much as 25. So that is all fine. I could also transform this to say that there might be a younger and an older group. I might have students. And I say, okay, if you are under 29 years old, then you belong to the younger students. And if you are older than 30 or exactly 30, then you belong to the older students.

That is an arbitrary decision, and then it would be the ordinal level. If I wanted to, although I am not sure whether that makes sense in this case, I could of course break it down even further, for example by age, or simply form groups of age categories that I do not want to order for some reason.

Okay, the point is that you can always transform it in one direction. You can always move from a higher level to a lower one, but of course this does not work in the other direction. This is important to recognize when you are collecting data and there is a concept that seems very important to you and that you want to use in your statistical analysis.

So try to collect the information at a very high level whenever possible.

And one final important point: I already mentioned that Likert-scale data are ordinal data, and that is of course correct. But in research, we often do not ask just one question—we ask several questions. These are known as scales or item batteries. And when you take the mean of this Likert-scale data, strictly speaking, you do not obtain interval data, but you can treat it as such.

The original article by Stevens (1946)

Stevens’s (1946) article “On the Theory of Scales of Measurement” addresses the fundamental theories and concepts of measurement scales and their applications in the sciences. Stevens defines measurement as the assignment of numerals to objects or events according to specific rules and classifies measurement scales into four main types, based on the underlying empirical operations and mathematical properties:

  1. Nominal scale:
    • Used to categorize or classify objects.
    • Numbers are used solely as labels, without representing an order or magnitude.
    • Applicable statistical methods: mode and frequency analysis.
  2. Ordinal scale:
    • Determines the order or ranking of objects.
    • Distances between values are not defined.
    • Examples: mineral hardness, intelligence ratings, and personality assessments.
    • Statistics such as the median and percentiles can be applied, but require caution.
  3. Interval scale:
    • Measures distances between values, with an arbitrary zero point.
    • Examples: temperature measured in Celsius or Fahrenheit.
    • Permissible statistical methods: mean, standard deviation, and correlations.
  4. Ratio scale:
    • Includes an absolute zero point and allows statements about ratios.
    • Examples: length, weight, and time.
    • All statistical operations are possible, including logarithmic transformations.

Stevens emphasizes that the choice of scale determines the types of statistical analyses that can be applied to the data. He argues that measurements are defined by clear rules and proposes using this definition as the basis for evaluating and classifying measurement scales.

The article concludes by stating that no scale is perfect, since measurements are always limited by the precision and accuracy of the underlying empirical operations. Nevertheless, these classifications provide a basis for the systematic investigation of measurement procedures across different scientific disciplines.