Skip to content

Data Catalog ​

The Data Catalog holds the datasets your dataflows read: the ones that come with Curio, files you import from your computer, and results your own nodes save. Add a dataset to a dataflow, drag it onto the canvas, and Curio writes the code that loads it.

Find a dataset ​

On the canvas, open Data > Data Catalog, or open the Data Catalog dropdown in the Tools panel on the left and click Browse Data Catalog +. The drawer has three tabs: Browse all, In project for the datasets this dataflow uses, and Computed for outputs your nodes saved. In the catalogs, a project is one of your saved dataflows.

Click Add to project on a dataset and confirm, and it appears in the Tools panel's Data Catalog dropdown. Remove from project takes it out again.

The Data Catalog tab at the top of the Projects page shows the same datasets for your whole account. Click one to open its panel, where Add to all projects adds it to every project you have and to new ones. When the server allows it, the panel also offers Publish, to share a dataset of yours with everyone on that server.

The Data Catalog drawer on the canvas, where a dataset is added to the dataflow and then appears in the Tools panel.

Load it on the canvas ​

Drag a dataset from the Tools panel onto empty canvas, and Curio creates a Data Loading node with code that reads it the right way for its format. Drop it onto an existing node, and the loading code is added to that node instead. The code names the dataset rather than a file path, so it keeps working when you share or move the dataflow.

Clicking a dataset in the dropdown highlights the nodes that use it. A node that reads a dataset shows a DATASET pill in its title bar; click the pill to find the dataset in the Tools panel.

A dataset dragged from the Tools panel onto the canvas becomes a Data Loading node, which is then connected to a transformation node and run.

Bring your own data ​

Click Import dataset at the bottom of the drawer, or at the top of the Data Catalog page, and pick a file. Curio accepts:

  • CSV (.csv), JSON (.json) and Parquet (.parquet) files
  • GeoJSON (.geojson) and Shapefiles (.shp)
  • GeoTIFF rasters (.tif, .tiff)
  • OpenStreetMap extracts (.pbf) and GeoPackages (.gpkg), which become one dataset per layer

Importing does not add the file to the open dataflow, so click Add to project on it next. To get a dataset from an open data portal, storage or a service instead, see Discovery Catalog.

Delete removes one of your uploads or computed datasets from your account and from every dataflow that uses it. Removing an upload from the last dataflow that uses it also deletes it, and the confirmation says so first.

Save node outputs as datasets ​

Next to a node's play button is a small database toggle, off by default. Turn it on and run the node: its output is saved to your Data Catalog as a computed dataset and appears in the drawer's Computed tab, ready to add to any dataflow. Running it again replaces the saved copy, and the dataflow must be saved for its outputs to be stored.

Tables are saved as Parquet (a GeoDataFrame keeps its geometry and coordinate system), rasters as GeoTIFF, plain Python values as JSON, and a tuple of results as a multi-part dataset. A node whose output was saved shows an OUTPUT pill.

Follow a dataset's lineage ​

View details on a dataset opens its Overview, Schema, Table Preview and Lineage tabs. For a computed dataset, Lineage shows the node it was Generated by, that node's inputs, and the nodes it is Consumed by, then lists the projects that use it. The record stays with the dataset even after you remove it from every dataflow. Export downloads most datasets as a file.

A dataset's details open on the Lineage tab, which shows the node in the saved dataflow that consumes it.