Datasets

A MIDAS project consists of data loaded from CSV files and derived data generated through SQL queries and other transformations. This page describes the different dataset types and how they behave.

Dataset Types

Primary Dataset

When you load a CSV or TSV file, a Primary Dataset is created. It stores the imported data as-is and supports direct cell editing and row exclusion.

Project Overview displays metadata such as the original file name, import date, and file size.

Derived Dataset

Derived Datasets are generated by transformation operations including The SQL Query Editor Tab, Crosstab, Reshape, and Dummy Coding, or by Save as Dataset in Filtered Data. Each Derived Dataset records the operation that produced it and any parent datasets it depends on.

You cannot directly edit data in a Derived Dataset. To change the data, modify the transformation operation or update the parent dataset.

An Ephemeral Dataset is a temporary dataset that is not persisted in the project. Unlike Primary and Derived Datasets, which are saved in the project, an Ephemeral Dataset exists only while its tab is open and is discarded when you close the tab. Filter results in the Filtered Data tab are an example.

If you save the project while the tab is still open, the data is included in the project file so the tab can be restored. Ephemeral Datasets not referenced by any open tab are excluded when you save. To keep a filter result permanently, save it as a Derived Dataset with Save as Dataset in Filtered Data or Save Filtered Data in Data Table.

Automatic Data Type Inference

When you load a CSV file, MIDAS automatically determines the data type for each column. See Data Preparation and Import for the supported data types (boolean, int64, float64, date, datetime, string, enum).

Inference follows a priority order: boolean is checked first, then numeric types (int64/float64), then date types (date/datetime). If none match, the column is assigned string type. The enum type is never auto-inferred -- create it by manually converting from string or int64 type.

Empty cells and missing values are treated as null and do not affect type inference.

Schema Inheritance in Derived Datasets

When you create a Derived Dataset via SQL Query Editor or other operations, the resulting columns inherit metadata from the parent datasets. Specifically, the measurement scale (nominal, ordinal, interval, ratio), data type, and enum name are inherited.

The inheritance rules are:

  • If a result column has the same name as a column in the parent dataset, the parent's settings are inherited
  • The measurement scale, data type, and enum name are inherited independently. When multiple parents exist (e.g., JOIN), the first dataset in the FROM clause takes priority for a column found in several parents, and an attribute missing from an earlier parent is filled in from a later one
  • If no matching parent column exists, the type is determined automatically from the query result

For example, if you changed a zip code column from interval to nominal scale in the parent dataset, running SELECT zip_code FROM ... in SQL Query Editor preserves the nominal scale setting.

However, when a transformation like CAST(category AS INTEGER) changes the result type, the parent's measurement scale is not inherited and is determined automatically from the result type instead. Adjust it manually if the automatic result does not match your intent.

Even when a column inherits an enum name, if the result data contains values outside the enum definition, the column becomes string type and the enum name is dropped. For example, this applies to a column where COALESCE or CASE substitutes a value not in the definition.

Cascade Deletion

Deleting a dataset cascades to dependent resources.

Deleted

  • Derived Datasets that depend on the deleted dataset (transitively, including grandchildren)
  • Models fitted on the deleted dataset

Not deleted

  • Reports (though report elements referencing deleted datasets are automatically removed from the report)

For example, if Primary Dataset "A" has Derived Dataset "B", and "B" has Derived Dataset "C" and Model "M", deleting "A" also deletes "B", "C", and "M".

To review the impact before deleting, check the dependency graph in Project Lineage.

Deleting a Model

Deleting a model cascades in the same way.

Deleted

  • Derived Datasets produced by the model (such as diagnostic datasets and coefficient tables)
  • Derived Datasets and Models that depend on those datasets (transitively)

Not deleted

  • The dataset the model was fitted on
  • Reports (though report elements referencing the deleted model or its associated datasets are automatically removed from the report)

For example, if Model "M" was fitted on Dataset "A" and produced coefficient dataset "M-coef", and "M-coef" has Derived Dataset "D", deleting "M" also deletes "M-coef" and "D", but "A" remains.

Lazy Evaluation and Caching

Derived Datasets do not compute their data until it is needed. For example, immediately after opening a project file, Derived Dataset data has not yet been computed. The computation runs when you first reference the dataset -- in a Data Table tab, Graph Builder, an analysis tab, Statistics, a Report, or a CSV export -- and the result is cached in memory. While the computation runs, Graph Builder and Reports show "Evaluating derived dataset...", and if it fails they show the reason in the same place.

When a parent dataset is updated (via reload, type change, cell edit, etc.), all downstream Derived Dataset caches are discarded. Data is automatically recomputed the next time it is needed.

Dependent Models are automatically re-estimated. Their estimation results are hidden until re-estimation completes. If re-estimation fails, a notification appears at the bottom-right of the screen. Use the Open Model button in the notification to open the model and check the reason for the failure.

Unsaved fit results in analysis tabs also become invalid when the dataset used for the fit is updated. Unlike saved Models they are not re-estimated automatically, so the GLM, GLMM, and Linear Regression tabs remove the results from the view and show a message asking you to run the analysis again.

Saving Computed Data with the Project

A Derived Dataset with Save data with project enabled saves its computed data in the project file (MDS). This setting is disabled by default. When it is disabled, the MDS file contains only the definition of the operation, and MIDAS re-runs the operation when the data is first needed after the project is opened.

Enabling it makes the MDS file larger, but the data is available as soon as the project is opened, without waiting for recomputation1. The difference grows with datasets produced by slow SQL queries or transformations of large data. Toggle the setting from the Datasets section in Project Overview or from the menu of the Data Table tab. The Agent API can also change it.

The setting decides only whether the data is saved; it does not change when the data is recomputed. Even when it is enabled, MIDAS discards the computed data in memory when a parent dataset is updated, and recomputes it the next time it is needed, as described in Lazy Evaluation and Caching. The contents of the MDS file change the next time you save the project.

The row data of a dataset with this setting enabled is included in the MDS file. Be mindful of that data when sharing the file with others.

Renaming Datasets and Automatic SQL Updates

When you rename a dataset in Project Overview, SQL queries in Derived Datasets that reference the old name are automatically updated.

For example, renaming "sales_2024" to "sales" changes FROM "sales_2024" to FROM "sales" in any dependent SQL. The update is based on SQL parsing, so occurrences in string literals or comments are not affected.

The update applies to table references enclosed in double quotes, such as FROM "sales_2024". References written without quotes, such as FROM sales_2024, are not updated -- edit the SQL manually after renaming.

After renaming, affected Derived Dataset caches are discarded and recomputed on next access.

The automatic update applies to saved SQL; the query shown in an open SQL Query Editor tab is not rewritten. See SQL Query Editor Tab for what happens when you rename a dataset while an edit tab is open.

See also

Footnotes

  1. The MDS file includes only data that has been computed at the time of saving. If you save the project after a parent update has discarded the computed data and before it is recomputed, the data of that dataset is not included in the MDS file, and MIDAS recomputes it when it is first needed after the project is opened. ↩