The Dummy Coding Tab
The Dummy Coding tab converts categorical variables (nominal/ordinal scale) into numeric dummy variables (0/1). Use this to prepare categorical variables as predictors for regression analysis or GLM.
Basic Usage
Opening Dummy Coding
Select Data > Dummy Coding... from the menu bar to open a new Dummy Coding tab.
Configuring the Transformation
- Select the target dataset from the Dataset dropdown
- Configure each column's Scale and Action, and optionally the Reference (reference category), in the column table
- Review each column's reference category and the number of dummy variables in the Encoding Preview section
- Click Create Dataset
- Enter a dataset name and click OK
In the Encoding Preview, check that the unique value count is what you expect, with no extras from typos or inconsistent labels, and that the reference category is the one you intend. The preview also reports the reference category's observation count as n =. A column with only one unique value cannot be encoded and shows an error.
Sample Data Used in This Page
The examples on this page use a survey dataset (survey.csv) for 5 people. blood_type and education are categorical variables.
| name | age | blood_type | education |
|---|---|---|---|
| Alice | 28 | A | college |
| Bob | 35 | B | graduate |
| Carol | 42 | O | college |
| Dave | 31 | A | high_school |
| Eve | 26 | AB | graduate |
Column Configuration
Configure the Scale, Action, and Reference for each column.
Scale
| Scale | Description | Dummy Coding |
|---|---|---|
| nominal | Nominal scale | Available |
| ordinal | Ordinal scale | Available |
| interval | Interval scale | Not available (used as numeric) |
| ratio | Ratio scale | Not available (used as numeric) |
On an Enum column, the Enum definition order (low → mid → high, for example) determines both the order in which dummy variables are generated and the default reference category. This is the same order the graph axes use. For non-Enum columns, see Category Order. To base the coding on a specific category, select it in Reference.
The initial Scale is the measurement scale set on the column. For columns without a scale, it is inferred from the data type: string and boolean columns become nominal, and numeric and date/time columns become interval. See Data Types and Measurement Scales for details. Change the Scale from the dropdown as needed. String and enum columns store their values as text, so only nominal and ordinal are offered for them. A code recorded as a number — an equipment or line number, for example — is inferred as interval, but changing its Scale to nominal makes it available for dummy coding. Its categories are then sorted in numeric order, and the default reference category is the smallest number.
Action
| Action | Description |
|---|---|
| Not included | Exclude from the output dataset |
| Include as-is | Keep column as-is in the output |
| Dummy code | Convert to dummy variables (original column excluded) |
| Dummy code (keep original) | Convert to dummy variables and keep the original column. Choose this to use the original categorical column for graphs or crosstabs in the output dataset |
Dummy code options are only available for categorical columns (nominal/ordinal). Boolean columns are excluded1.
Reference
The reference category is the group that dummy variable coefficients are compared against. Choosing a meaningful baseline — a control group, a standard condition, or a pre-intervention level — makes the coefficients easier to interpret.
For columns with the Action set to Dummy code or Dummy code (keep original), select the reference category from the Reference dropdown. The default is the first category in the category order.
If the selected reference category is no longer present in the data, the Dummy Coding tab shows an error before the transformation runs. When rows with that category disappear from the source dataset, for example after Reload Dataset, an error message appears and the Reference dropdown marks the value with a not in data note. Create Dataset / Update Dataset stay disabled until another reference category is selected.
The Reference dropdown reports each category's observation count as n =. The reference category's observation count affects the standard error of every dummy variable coefficient generated from that column. In a linear regression whose only predictor is that categorical variable, choosing a category with few observations as the reference widens the interval estimates of all coefficients for that column2. In a model with other predictors, the coefficient variances also carry terms from the correlations among the predictors, so the reference category's observation count alone does not determine them.
The counts shown here are the non-missing rows of that column alone. A regression or GLM applies listwise deletion to rows missing on any column the model uses, the response included, so each category holds fewer rows than shown here whenever such rows exist.
Columns With Many Categories
Setting a column with more than 50 unique values to Dummy code makes the Dummy Coding tab show a warning that reports how many dummy variables will be generated. Because k categories produce k-1 dummy variables, setting a column that holds a distinct value per row, such as a name or an ID, brings the output dataset's column count close to its row count.
The warning does not stop the transformation. If the setting is intentional, Create Dataset runs as usual. The transformation time and the size of the output dataset both grow with the number of categories.
For columns with more than 50 unique values, the Reference dropdown lists only the first 50 categories in category order and reports how many are left out. Expanding tens of thousands of options would stall the settings screen. To use a category outside that range as the reference, specify it through referenceCategories in the Agent API.
How It Works
From k categories, k-1 dummy variables are generated, omitting one category as the reference category3. This scheme is called treatment coding (coding relative to a reference category).
- Extract unique category values4
- Sort them in the category order
- Exclude the reference category. The default is the first category in the sort order, which can be changed from the Reference dropdown
- Generate a dummy variable for each of the remaining k-1 categories
- Code 1 for matching rows, 0 otherwise
Category Order
The category order determines the order in which dummy variables are generated and the default reference category. It is the same order graph axes and legends use, and it is derived from the column type. For Enum columns, categories follow the Enum definition order regardless of the measurement scale. For int64 and float64 columns, categories are sorted in numeric order. For all other columns, categories are sorted by the Unicode code point order of the category names. Code point order matches dictionary order for names that only use letters and digits of the same case, but all uppercase letters (A–Z) sort before all lowercase letters (a–z). The order does not depend on the language settings of the environment MIDAS runs in.
Example
Converting the blood_type column (unique values: A, AB, B, O):
- Reference category:
A(the default, first in the category order) - Dummy variables generated:
blood_type_AB,blood_type_B,blood_type_O
| blood_type | blood_type_AB | blood_type_B | blood_type_O |
|---|---|---|---|
| A | 0 | 0 | 0 |
| B | 0 | 1 | 0 |
| O | 0 | 0 | 1 |
| A | 0 | 0 | 0 |
| AB | 1 | 0 | 0 |
Rows with the reference category A have all dummy variables set to 0.
When the dummy variables are used in a regression model, each dummy variable coefficient is an estimate of the difference from the reference category. The scale on which that difference is measured depends on the model's link function. In linear regression (identity link), it is a difference on the scale of the response variable itself. In GLMs with a non-identity link — logistic regression, Poisson regression — it is a difference on the scale of the linear predictor, such as log-odds or a log rate. In this example, fitting blood_type_B in a linear regression gives a coefficient that estimates the difference in the response variable between type B and the reference type A. In a model with other predictors, this is the difference holding them constant.
Output Dataset
With the sample data above, setting name to Not included, age to Include as-is, and both blood_type and education to Dummy code produces the following output.
| age | blood_type_AB | blood_type_B | blood_type_O | education_graduate | education_high_school |
|---|---|---|---|---|---|
| 28 | 0 | 0 | 0 | 0 | 0 |
| 35 | 0 | 1 | 0 | 1 | 0 |
| 42 | 0 | 0 | 1 | 0 | 0 |
| 31 | 0 | 0 | 0 | 0 | 1 |
| 26 | 1 | 0 | 0 | 1 | 0 |
blood_type (4 categories) produces 3 dummy variables, and education (3 categories: college, graduate, high_school; reference: college) produces 2. Column names follow the {original_column_name}_{category_name} format, with data type int64 (0 or 1) and a ratio measurement scale. They go straight into regression or GLM as predictors and are not dummy coded again. Row count stays the same. The original dataset is not modified; results are saved as a new derived dataset.
Editing Saved Dummy Coding
Existing Dummy Coding datasets can be edited. Right-click a dataset in the Project Lineage tab, or open the dataset menu (⋮) in Project Overview, and select Edit Operation... to open Dummy Coding in edit mode with the source dataset, column actions, reference categories, and scale overrides restored.
In edit mode, if other datasets, models, or reports depend on this dataset, a warning is displayed showing the type and count of affected items. Changing the source dataset carries over the settings of columns that exist in the new source under the same name. Actions of columns that are no longer categorical change to Include as-is, and reference categories that cannot be confirmed in the new source are reset. Settings that could not be carried over are listed in a notice. Click Update Dataset to apply the changes. Derived datasets that depend on this dataset are recalculated the next time their data is needed, and dependent models are automatically re-estimated.
If the transformation is changed from another path such as another edit tab or the Agent API between the time this tab opens and the time you click Update Dataset, MIDAS stops the save and shows a dialog. This keeps the earlier change from being lost without warning. The dialog shows the transformation as it was when the tab opened next to the version changed elsewhere since then. Choose Overwrite with This Tab's Edits to save the edits in this tab over the version changed elsewhere. Choose Keep Editing to return to the edit tab without saving; the edits stay in place, and saving again shows the same dialog.
Next steps
- Linear Regression - Regression analysis using dummy variables
- The GLM Tab - GLM using dummy variables
See also
- The Convert Column Types Tab - Converting data types
Footnotes
-
Boolean columns are already equivalent to 0/1, so select Include as-is to use them directly. ↩
-
In a linear regression whose only predictor is that categorical variable and whose error variance σ² is the same in every category, the variance of the coefficient for category j is σ²(1/n_j + 1/n_ref). σ² here is the variance of the error conditional on the category, not the variance of the response itself. n_j is the number of observations in category j and n_ref the number in the reference category; the 1/n_ref term is shared by every category's coefficient. ↩
-
Creating dummy variables for all k categories would make their sum equal 1 in every row. The intercept is also 1 in every row, so the two would be linearly dependent and the regression coefficients could not be estimated uniquely. Omitting one category avoids this problem. ↩
-
Missing values are not counted as unique values. Rows where the original column is missing have missing values in all generated dummy variables and are dropped by listwise deletion in later regression or GLM analyses. Columns with only one unique value cannot be dummy coded. ↩
Also available as a Markdown file.