Data Preparation and Import

Loading a File

Click the Open File button on the launcher screen and select a file, or drag and drop a file onto the launcher screen. When you drop multiple CSV or TSV files at once, each file is imported as a separate dataset in one project. Drop MDS and ZIP files one at a time. The Open from URL button loads a CSV, TSV, or MDS file from a URL. To use sample data, choose from the "Sample Data" section on the launcher screen. The New Empty Project button opens a project with no datasets. See Getting Started for detailed steps.

When MIDAS reads a CSV or TSV file from the launcher screen with warnings, it lists them before opening the project. The warnings cover a Content-Type in the URL response that names a media type other than CSV, TSV, or plain text, a Content-Type charset that does not match the data, and renamed column names. Open anyway opens the project. Back returns to the launcher screen without creating a project. When the file comes from Open from URL, Back returns to that dialog with the URL, encoding, and delimiter kept, so you can select another encoding or delimiter and open the file again. Warnings about URLs not in the trusted URL list and about file size are confirmed before or during the download and are not listed again.

After opening a project, you can load additional datasets from the Data > Import Data... menu. In the Import Data dialog, select a local CSV or TSV file, or switch to the URL tab to fetch a CSV or TSV file from a URL. Fetching from a URL requires the server to allow CORS (Cross-Origin Resource Sharing). After you select a file or fetch a URL, the dialog shows the read settings (encoding, delimiter, and whether the first row is a header) and a preview. If the file cannot be read, the dialog shows the error on the same screen, so you can change the settings and read it again. Back returns to the screen for choosing a file or URL.

Supported File Formats

MIDAS supports four file formats: CSV, TSV, MDS, and ZIP.

CSV (Comma-Separated Values) The most common data format. Columns are separated by commas (,). File extension is typically .csv.

TSV (Tab-Separated Values) A file format where columns are separated by tab characters. File extension is .tsv or .txt.

MDS (MIDAS Project File) MIDAS's native project file format. Contains datasets, analysis settings, and reports. See Project File (MDS) for details.

ZIP (Multiple CSV/TSV Files) A ZIP archive containing CSV or TSV files. Each file is imported as a separate dataset.

Excel files (.xlsx) cannot be loaded directly. Save your spreadsheet as CSV from Excel's "Save As" menu.

Character Encoding UTF-8, Shift-JIS, and EUC-JP encodings are supported. Encoding is auto-detected, including for CSV and TSV files inside a ZIP archive. When loading from a URL, the charset declared in the server's Content-Type header is also used for detection. When the file contains bytes that cannot be decoded with the encoding used for reading (auto-detected or specified), MIDAS does not load the file and shows an error. If the detected encoding is not correct, you can specify the encoding in the Import Data dialog preview and in the Open from URL dialog. ZIP import has no encoding selector, so extract the archive and import the files individually when the detected encoding is wrong. When saving CSV from Excel, UTF-8 is recommended: select "CSV UTF-8 (Comma delimited)" format.

File Structure

MIDAS treats the first row as a header row. The values in the first row become column names, and subsequent rows become data. If your CSV does not have a header row, uncheck the "First row is header" checkbox in the Import Data dialog preview. MIDAS then generates column names automatically (Column1, Column2, ...) and treats the first row as data.

If the header row is empty or contains only blank cells, MIDAS rejects the file with an error instead of silently dropping data. Fix the file in a text editor and retry the import.

You can choose the delimiter from comma, tab, semicolon, and vertical bar (|). The default delimiter follows the file name extension: tab for .tsv and .txt, comma otherwise. MIDAS does not infer the delimiter from the file contents. To read a file whose delimiter differs from the default, specify the delimiter on the preview screen of the Import Data dialog or in the Open from URL dialog. Files dropped on the launcher screen, files opened with Open File, and files in a ZIP archive are read with the default delimiter. To read such a file with another delimiter, open a project and import the file from the Import Data dialog. If the specified delimiter does not match the file, the data is read as a single column or is rejected because of a column count mismatch.

If even one row has a different number of columns than the header (the first row when there is no header row), MIDAS rejects the import with an error so that data is not silently lost1. A row with a quoted value that is not closed correctly, and a row longer than 64 MB, are also rejected with an error. A quoted value must end with a quotation mark (") followed by a delimiter or a line break. For these errors, the error message shows the row number of the first problem found and the content of that row. Row numbers are counted from the top of the file, including the header row and blank lines2. A file that mixes line break types (CR, LF, CRLF) is also rejected, but that error message does not show a row number. Fix the file in a text editor and import it again.

Example:

Name,Age,Country
Alice,25,USA
Bob,30,Japan
Charlie,28,UK

Missing Values Empty cells in CSV files are loaded as missing values (null). Strings such as NA, ., and - are not treated as missing values; they are loaded as text. See Data Types and Measurement Scales for details. Missing values are excluded from statistical calculations and graph rendering. Rows containing missing values are not removed from the dataset3.

Data Types

MIDAS automatically determines data types when loading. The supported data types are boolean, int64, float64, date, datetime, string, and enum. See Data Types and Measurement Scales for the details of each type.

When you load a CSV file, MIDAS creates two datasets: one that keeps the file values as text, and one converted to the detected types. The text dataset is named with a (raw) suffix, and the converted dataset inherits the original name. The converted dataset opens by default and is the one to analyze. If every column is detected as text, only one dataset is created.

Data types are displayed below the column name as int64. If a detected type is not what you intended, adjust the conversion with The Convert Column Types Tab. Because the raw dataset keeps the original file values, changing the target type does not lose information. Cell editing, row exclusion, and row comments are done on the raw dataset; changes to the raw dataset are automatically reflected in the converted dataset.

Measurement Scales

MIDAS automatically assigns a measurement scale (Nominal, Ordinal, Interval, Ratio) to each column. Measurement scales affect the available graph types and statistical methods. Right-click a column in the Data Table and change its scale from Edit Scale of Measurement.

See Data Types and Measurement Scales for what each scale means and how it affects analysis.

Next steps

See also

Footnotes

  1. A row whose extra columns are only empty values at the end (such as 1,2, in a two-column file) is not rejected; MIDAS reads it without the trailing empty values. ↩

  2. When a quoted value contains line breaks and spans several lines, the row that contains it counts as one row. In that case, the row number in the error message is smaller than the line number in a text editor. ↩

  3. In a single-column CSV, MIDAS skips a row whose value is empty (including a row that contains only "" and a row that contains only delimiters) in the same way as a blank line. ↩