Data Management in SPSS

Introduction

In this article, I will guide you step by step through the world of data management in SPSS. While R is popular among many researchers, SPSS (Statistical Package for the Social Sciences) has been one of the most widely used programs for statistical analyses for decades—especially in the social and behavioral sciences. Perhaps you have opened SPSS before, clicked around a bit, and then wondered where best to start.

Here, we will look at how to import data into SPSS, prepare, filter, and code it, handle missing values, and then export your “clean” data again. We will examine the classic “click-and-menu” approach on the one hand, while also learning a little syntax on the other. As in R, this makes reproducibility easier.

Why use SPSS for data management?

You may be wondering: “Why should I choose—or continue using—SPSS?” Well, of course, it is a matter of personal preference. SPSS’s main advantage is that it has a graphical user interface (GUI) that feels intuitive to many people. You can click through menus, set filters, calculate new variables, or define missing values—all without having to learn a programming language.

However, SPSS also has a syntax language that allows you to retrace and document every step precisely, so you can easily repeat your analysis later. This is a major advantage if you want to work reproducibly.

SPSS is also already licensed at many institutions (universities, research departments, etc.). So you often do not have to pay anything extra because your institution covers the cost.

SPSS: Understanding the Interface and Setup

The SPSS Views

SPSS has two main views that you can switch between:

  1. Data View: This is where you see the rows (cases) and columns (variables).
  2. Variable View: This is where you define the properties of each variable (name, data type, label, possible values for categorical variables, etc.).

When you start SPSS, you usually land in Data View. Simply switch between “Data View” and “Variable View” at the bottom left.

A Variable Label is a more detailed description of your variable (e.g., “Age of respondents” instead of just “age”). Value Labels are specific descriptions for numerical codes representing answers, such as 1 = “male” and 2 = “female”.

Syntax and Output

  • Syntax window: If you select “File -> New -> Syntax” from the menu at the top, a window opens where you can enter SPSS commands, such as FREQUENCIES, COMPUTE, etc.
  • Output window: Results are displayed in a separate window called the “Output Viewer.” There you will find tables, charts, and error messages.

Many people start using SPSS as a point-and-click program, but especially for data management, it is worthwhile to use the syntax from time to time. This saves you time when you want to carry out similar steps repeatedly.

Importing Data into SPSS

Opening SPSS Files (.sav)

If you have a .sav file (the native SPSS format), you can simply open it by double-clicking it (if your system is configured accordingly) or via the menu File -> Open -> Data….

Importing CSV/TXT Files

If you have CSV files (Comma-Separated Values), proceed roughly as follows:

  1. File -> Open -> Data…
  2. At the bottom, under “Files of Type,” select “Text (.txt, .dat, .csv).”
  3. Locate your CSV file.
  4. The Text Import Wizard starts. There, you specify whether you have headers (i.e., column names in the first row), which delimiter to use, etc.

Syntax example for a CSV file:

GET DATA
  /TYPE=TXT
  /FILE='C:\meinprojekt\rohdaten\umfrage.csv'
  /DELCASE=LINE
  /DELIMITERS=","
  /QUALIFIER='"'
  /ARRANGEMENT=DELIMITED
  /FIRSTCASE=2
  /VARIABLES=
   ...
  .
EXECUTE.

The VARIABLES sections specify how you want to name columns and which formats to use. With the “Text Import Wizard,” you can have a syntax block generated automatically.

Loading Excel files into SPSS

For .xlsx or .xls:

  1. File -> Open -> Data…
  2. Select “Excel (.xls, .xlsx)” as the file format.
  3. In the wizard, you can specify whether the first row contains column names.

The syntax example is considerably more extensive than in R—usually, it is generated via the wizard. A shortened version could look like this:

GET DATA
 /TYPE=XLSX
 /FILE='C:\meinprojekt\rohdaten\umfrage.xlsx'
 /SHEET=name 'Tabelle1'
 /CELLRANGE=FULL
 /READNAMES=ON
 /ASSUMEDSTRWIDTH=32767.

Afterward, don’t forget EXECUTE. so that SPSS actually loads the data.

Data preparation and transformation

The raw data is now in SPSS, but you will usually need to make some adjustments before you can start your analyses.

Creating, recoding, and otherwise modifying variables

Computing a new variable (COMPUTE)

Suppose you have two variables, groesse (in m) and gewicht (in kg), and want to calculate the BMI. You can do this via:

Menu:

  • Transform -> Compute Variable…

Name of the new variable: bmi
Formula: bmi = gewicht / (groesse ** 2)

Clicking OK creates the variable and adds it as a new column in the Data View.

Syntax:

COMPUTE bmi = gewicht / (groesse ** 2).
EXECUTE.

Recoding values (RECODE)

For example, if you want to create the categories “low,” “medium,” and “high” from a variable einkommen, you can tell SPSS: everything below 1500 = low, 1500–3000 = medium, and 3000 or more = high.

Menu: Transform -> Recode into Different Variables

Syntax (example):

RECODE einkommen (LO THRU 1499 = 1) (1500 THRU 2999 = 2) (3000 THRU HI = 3)
  INTO einkommen_gruppe.
EXECUTE.

VALUE LABELS einkommen_gruppe
  1 "Gering"
  2 "Mittel"
  3 "Hoch".
EXECUTE.

So, now einkommen_gruppe is a new variable with values 1, 2, and 3. The VALUE LABELS make them more readable.

“Recoding” means translating certain ranges of values or categories into new values. This is similar to R, where you might use if_else().

Filtering and selecting cases

  1. Select Cases: Menu -> Data -> Select Cases…
    • Examples: alter > 30 (only people older than 30), geschlecht = 1 (only men, if 1 is coded as male).
    • SPSS then hides or deletes all other rows, depending on your selection.
  2. Split File: Menu -> Data -> Split File…
    • This allows you to conduct separate analyses for each group, for example, separately by gender.

Syntax:

USE ALL.
COMPUTE filter_$=(alter>30).
FILTER BY filter_$.
EXECUTE.

You have now filtered all people with alter>30. Don’t forget to turn off the filter later (FILTER OFF).

Sorting

Menu: Data -> Sort Cases.
Syntax:

SORT CASES BY einkommen (D).

(D) stands for descending. (A) would stand for ascending.

Handling missing values

What are missing values in SPSS?

In SPSS, missing values are usually displayed as a dot . (similar to NA in R). You can also define user-defined missing values, such as -99 for “no response.”

Declaring missing values

In the Variable View, you can enter under “Missing” which values SPSS should interpret as missing. This can be, for example, -99 or a range from -99 to -90. SPSS then excludes these values from the statistics.

Syntax (example):

MISSING VALUES einkommen (-99).

Checking frequencies

Do you want to know how many NAs or missing values there are?

Menu: Analyze -> Descriptive Statistics -> Frequencies…
Syntax:

FREQUENCIES VARIABLES=einkommen
 /FORMAT=NOTABLE
 /STATISTICS=MIN MAX MEAN
 /MISSING=REPORT.

This gives you a small table showing, among other things, how many missing values exist.

Strategies for dealing with missing values

  • Listwise Deletion: Rows with missing values in any relevant variable are excluded completely.
  • Pairwise Deletion: For some statistics, SPSS calculates using different “subsamples” – this can lead to fluctuations in the number of cases.
  • Imputation: In SPSS, you can use the Missing Values Analysis (MVA) module to perform multiple imputation, for example. This is a more advanced method that we probably won’t look at here.

Data export and storage

Saving in SPSS format (.sav)

File -> Save As… – Select Save as type: SPSS Statistics (*.sav). This way, you have your data in a format that allows you to preserve all formatting and label information.

Export as CSV

If you want to share your data externally or import it into another tool, you can do the following in SPSS:

Menu: File -> Export -> [Optionen variieren je nach SPSS-Version], or:

SAVE TRANSLATE
  /OUTFILE="C:\meinprojekt\output\export.csv"
  /TYPE=CSV
  /MAP
  /REPLACE
  /FIELDNAMES.
EXECUTE.
  • /TYPE=CSV specifies the format.
  • /FIELDNAMES ensures that the variable names appear in the first row.
  • /MAP provides an excerpt from the log showing which variables were exported and how.

Exporting Excel files

Menu: File -> Export -> … (depending on the SPSS version, this may be done via “Save As…” -> Type: Excel file)
Syntax: This is relatively complex, similar to GET DATA for Excel. It is often easier to use the GUI.

Syntax files

To save not only the data but also your work steps, it is a good idea to:

  • Copy everything you do by clicking through the menus into a syntax file using “Paste”, or
  • Work directly in a syntax file.

This gives you a script-like workflow that you can reproduce later – similar to R scripts.

Best Practices & Tips

Folder Structure

Even though SPSS does not have a native “project” concept like RStudio, it is a good idea to have a clear folder structure. For example:

  • raw_data/ for raw data
  • processed_data/ for cleaned versions
  • syntax/ for all your syntax files
  • output/ for results, tables, and exported files

Documentation

  • Variable Labels and Value Labels are worth their weight in gold. You’ll be glad later that you don’t have to puzzle over what “var43” actually was.
  • If you use user-defined missing values, be sure to document why you define certain values as missing.

Syntax vs. Clicking

  • If you are new to SPSS and find writing syntax difficult, start with the menus. Then click “Paste” to see which syntax SPSS generates.
  • Once you feel more comfortable, use syntax right away – it saves time and frustration, especially with repetitive tasks.

A Word on Performance

SPSS is not primarily designed for huge datasets (Big Data). It can become slow when working with millions of cases. In that situation, it may be worth considering other tools (SQL, R, Python) or the SPSS Server module. But for medium-sized datasets in psychology, sociology, medicine, and similar fields, SPSS is usually more than sufficient.

A Quick Data Check

Before you start with in-depth analyses, take a look at frequencies and descriptive statistics. For example:

  • Descriptive Statistics -> ANALYZE -> DESCRIPTIVE STATISTICS -> DESCRIPTIVES…
  • Explore -> ANALYZE -> DESCRIPTIVE STATISTICS -> EXPLORE… (boxplots, outliers, etc.)
  • FREQUENCIES for categorical data.

This allows you to quickly identify irregularities or outliers.

A Practical Example Step by Step

Below, I’ll show you a short practical example (simplified) that demonstrates the entire workflow:

  1. Data import: Imagine you have a CSV named survey.csv with the columns id, gender, age, and income.
  2. Data preparation:
    • Divide age into age groups
    • Define missing values
    • Sort by income
    • If necessary, apply a filter to specific groups
  3. Export as an SPSS file

Import (syntax)

GET DATA
  /TYPE=TXT
  /FILE='C:\meinprojekt\raw_data\umfrage.csv'
  /DELCASE=LINE
  /DELIMITERS=","
  /QUALIFIER='"'
  /ARRANGEMENT=DELIMITED
  /FIRSTCASE=2
  /VARIABLES=
    id F1.0
    geschlecht F1.0
    alter F3.0
    einkommen F8.2
  .
EXECUTE.

Afterwards, you should see your four columns in Data View. You can use Variable View to adjust variable names, add value labels, and so on.

Creating age groups

RECODE alter
  (LO THRU 17=1)
  (18 THRU 64=2)
  (65 THRU HI=3)
  INTO alter_grp.
EXECUTE.

VALUE LABELS alter_grp
  1 "Kind/Jugend"
  2 "Erwachsen"
  3 "Senior".
EXECUTE.

Missing values

Assume that -99 represents “no information” in the income field. Then:

MISSING VALUES einkommen (-99).

Sorting and filtering

SORT CASES BY einkommen (D).  /* Absteigend nach Einkommen sortieren */

COMPUTE filter_$=(geschlecht=1).  /* Nur geschlecht=1, z.B. männlich */
FILTER BY filter_$.
EXECUTE.

Exporting the cleaned data

SAVE OUTFILE='C:\meinprojekt\output\umfrage_bereinigt.sav'
  /COMPRESSED.
EXECUTE.

You now have your “final version” in SPSS format.

Questions & Answers (Q&A)

To wrap up, here is a short set of questions so you can test your knowledge. For each one, I provide a model solution as an example answer.

  1. Question: How do I open a CSV file in SPSS without clicking through everything in the menu?
    Answer (model solution): You can open the Import Wizard from the menu and click “Paste” there. This generates a GET DATA command in the syntax. You then only need to append EXECUTE., and you have your syntax block. Manually, it could look like this:GET DATA /TYPE=TXT /FILE='C:\pfad\zu\datei.csv' /DELCASE=LINE /DELIMITERS="," /FIRSTCASE=2 /VARIABLES= ... . EXECUTE.
  2. Question: How can I define missing values if, for example, “-99” indicates “no answer”?
    Answer (model solution): In Variable View, I select “Missing” -> “Discrete values” -> “-99”. Or I write the following in the syntax:MISSING VALUES meine_variable (-99).
  3. Question: How do I create a new variable bmi when I have height and weight?
    Answer (sample solution): Go to “Transform -> Compute Variable…”, name the target variable bmi, and enter the formula = gewicht / (groesse ** 2). Using syntax:COMPUTE bmi = gewicht / (groesse ** 2). EXECUTE.
  4. Question: How do I create categories such as “low”, “medium”, and “high” from a continuous variable (e.g., income)?
    Answer (sample solution): Use “Transform -> Recode into Different Variables”. Syntax:RECODE einkommen (LO THRU 1499=1) (1500 THRU 2999=2) (3000 THRU HI=3) INTO einkommen_kat. VALUE LABELS einkommen_kat 1 "Low" 2 "Medium" 3 "High". EXECUTE.
  5. Question: Can I also share my final cleaned data with colleagues who do not have SPSS?
    Answer (sample solution): Sure, export it as a CSV or Excel file. For example, you can do this using:SAVE TRANSLATE /OUTFILE="C:\meinprojekt\output\export.csv" /TYPE=CSV /FIELDNAMES. EXECUTE. Or you can export it to Excel via the menu (File -> Export).
  6. Question: How can I document my analysis steps so that I can reproduce them exactly later?
    Answer (sample solution): Use syntax. Every click can be transferred to the syntax editor using “Paste”. Then save the .sps file. It contains all the commands, and you only need to run EXECUTE. This makes reproducible analyses and documentation much easier.
  7. Question: What should I do if I accidentally set Filter On but want to see all cases again?
    Answer (sample solution): Simply enter FILTER OFF in the syntax window and type EXECUTE. — or use the menu: Data -> Select Cases -> All cases -> OK. You will then have the full dataset available again.
  8. Question: How can I quickly check whether a variable contains outliers or unusual values?
    Answer (sample solution): I use “Analyze -> Descriptive Statistics -> Explore…” and take a look at the boxplots. Or simply Frequencies (Analyze -> Descriptive -> Frequencies). This often reveals whether, for example, a value such as “9999” occurs that should actually be defined as missing.

Summary

SPSS makes it easy to access your data, transform it, and ultimately export it as an SPSS file, CSV, or Excel file with just a few clicks. Nevertheless, you can use SPSS syntax to establish a reproducible workflow:

  1. Data import: Use GET DATA to load CSV, Excel, or other formats.
  2. Data preparation: Use Compute commands (or the “Compute” menu), Recode, Sort Cases, Select Cases, and define Missing Values.
  3. Missing values: Whether you use listwise deletion, pairwise deletion, or imputation depends on your analyses.
  4. Filtering & coding: Use Select Cases to filter, and RECODE to create new categories.
  5. Data export: Use SAVE OUTFILE='...' for .sav files and SAVE TRANSLATE ... TYPE=CSV for CSV files—depending on your needs.
  6. Documentation: Syntax is your friend. Copy everything (paste it) or write it directly into the syntax editor so that you can still understand your steps three months from now.

Advantages of SPSS:

  • Intuitive menus for beginners
  • A consistent interface that many students are familiar with (especially in the social sciences)
  • Syntax for reproducibility

Disadvantages or limitations:

  • Complex scripting logic is often less flexible than in programming languages such as R or Python
  • Not primarily designed for big data
  • License costs (unless covered by your university or organization)

Q&A (Sample questions from learners and brief model answers)

  1. Q: I have 20 columns containing items that all use the same response format (e.g., 1 = agree, 2 = disagree). Do I have to set this up separately for each column?
    A: In SPSS, you can use the “Copy-Paste function” in Variable View or enter VALUE LABELS var1 var2 var3 ... 1 "Agreement" 2 "Disagreement". in the syntax. This lets you define labels for multiple variables at once.
  2. Q: How can I “merge” different datasets in SPSS, for example if I have demographic data in one file and questionnaire data in another, and both files have an ID column?
    A: Menu: “Data -> Merge Files -> Add Variables…”. In the wizard, specify which ID variable should be used as the matching key. In syntax: MATCH FILES /FILE='C:\\path\\demographics.sav' /FILE='C:\\path\\questionnaire.sav' /BY id. SAVE OUTFILE='C:\\path\\merged.sav'. EXECUTE.
  3. Q: Why does SPSS always show 10 fewer cases than expected when calculating sums?
    A: Values are probably missing for 10 cases in one of the variables involved, and you have Listwise Deletion or MISSING=EXCLUDE enabled. You may want to use Pairwise Deletion instead or perform imputation.
  4. Q: I want to import my SPSS data into R. Is there a recommended approach?
    A: Yes, in R you can use the haven package: library(haven) and then read_sav("yourfile.sav"). This largely preserves value labels, even as factor levels.
  5. Q: Does SPSS have anything like loops or conditional execution if, for example, I want to generate syntax?
    A: Theoretically, yes, via the syntax features “DO IF… ELSE IF… END IF” and “DO REPEAT.” However, these are not as sophisticated as in traditional programming languages. For more complex data management scripts, R, Python, or Stata are more flexible.