Skip to content

Datasets (mtb.data)

Download the datasets and precomputed outputs the tutorials use into mtb.config.DEFAULT.data_path, or into the data_path= you pass.

Name Summary
mtb.data.fetch Download reference datasets into the data root, skipping those present.
mtb.data.fetch_outputs Download precomputed run_all outputs for a tutorial dataset.
mtb.data.fetchable Dataset ids mtb.data.fetch can download.

mtb.data.fetch

fetch(*datasets: str, data_path=None, quiet: bool = False) -> Path

Download reference datasets into the data root, skipping those present.

PARAMETERS DESCRIPTION
*datasets

Dataset ids to fetch; mtb.data.fetchable() lists them.

type str default ()

data_path

Data root that holds the dataset folders; None = mtb.config.DEFAULT.data_path. Each dataset lands in <data_path>/<dataset>/.

type Path | str | None default None

quiet

Suppress the one "downloading ..." line per dataset.

type bool default False

RETURNS DESCRIPTION
Path

The data root (not the dataset folder) - pass it as data_path=.

RAISES DESCRIPTION
ValueError

A dataset id is not a release asset (the message lists the ids).

RuntimeError

The downloaded archive lacks the <dataset>/ folder.

OSError

The download failed; the message names the URL.

Examples

>>> import multibench as mtb
>>> # the data root; D11 lands in <data_path>/D11/ (about 11 MB)
>>> root = mtb.data.fetch("D11")
>>> mtb.scan("D11", "vertical", data_path=root)
>>> mtb.data.fetch("D11", "D28", data_path="data", quiet=True)
Notes

What it covers. The tutorial reference sets published as release assets (mtb.data.fetchable()); the progress line states each one's approximate download size. The full collection is linked from the scMultiBench README ("Get the data" in the installation guide).

Idempotence. A non-empty <data_path>/<dataset>/ counts as present and is left alone, whatever it holds (even for an id that is not a release asset); an empty leftover folder is removed and fetched again.

Atomic extraction. Each archive is unpacked into a temporary folder under the data root and then moved into place. An archive entry that would write outside that folder raises RuntimeError.

See Also

mtb.data.fetchable : the dataset ids this function can download. mtb.data.fetch_outputs : precomputed run_all outputs for a tutorial dataset. mtb.config.Config : data_path, the default data root.

mtb.data.fetch_outputs

fetch_outputs(
    dataset: str, methods=None, *, data_path=None, quiet: bool = False
) -> Path

Download precomputed run_all outputs for a tutorial dataset.

Use it on a host without the method environments, such as Colab or a laptop. evaluate and plot then run on the stored outputs.

PARAMETERS DESCRIPTION
dataset

Dataset id with shipped outputs: D11, D28, D46 or D52.

type str

methods

Method ids that must be in the tree; checked only, the whole tree is downloaded either way.

type list of str default None

data_path

Data root that holds the dataset folders; None = mtb.config.DEFAULT.data_path.

type Path | str | None default None

quiet

Suppress the one "downloading ..." line.

type bool default False

RETURNS DESCRIPTION
Path

<data_path>/outputs/<dataset> - pass it to mtb.load_batch.

RAISES DESCRIPTION
ValueError

dataset has no shipped outputs (the message lists the ids).

KeyError

A name in methods is not in the tree (the message lists its methods).

RuntimeError

The archive lacks batch_result.json, or a non-empty outputs/<dataset>/ lacks one.

OSError

The download failed; the message names the URL.

Examples

>>> import multibench as mtb
>>> out = mtb.data.fetch_outputs("D11")
>>> res = mtb.load_batch(out)
>>> res.summary[["method", "status", "ARI", "NMI"]]
Notes

Tree contents. Exactly what mtb.run_all writes and mtb.load_batch reads: batch_result.json, long.csv, summary.csv and one folder per method holding its embedding.h5.

Selecting methods. methods only checks the names; the whole tree is downloaded. To restrict what is loaded, use mtb.load_batch(out, methods=[...]).

Idempotence. When <data_path>/outputs/<dataset>/batch_result.json exists nothing is downloaded. An empty leftover folder is replaced. A non-empty one without batch_result.json raises RuntimeError. Remove it and call again.

Atomic extraction. The archive is unpacked into a temporary folder and then moved into place. The archive may be rooted at the tree itself or at one folder (outputs-D11/, D11/); an entry that would write outside the temporary folder raises RuntimeError.

See Also

mtb.load_batch : load the downloaded tree as a BatchResult. mtb.run_all : produces the same tree from your own run. mtb.data.fetch : download the input datasets themselves.

mtb.data.fetchable

fetchable() -> list[str]

Dataset ids mtb.data.fetch can download.

The companion of mtb.available_datasets, which lists the ids with stored results; the two sets overlap but are not the same.

RETURNS DESCRIPTION
list of str

Dataset ids in natural order (D11, D28, D45, ...).

Examples

>>> import multibench as mtb
>>> mtb.data.fetchable()
['D11', 'D28', 'D45', 'D46', 'D52']
Notes

Overlap with stored results: D24 has published metric tables and no downloadable file, D46 downloads and has no stored table. The list is exactly the set of ids fetch() accepts.