Datasets (mtb.data)¶
Download the datasets and precomputed outputs the tutorials use into
mtb.config.DEFAULT.data_path, or into the data_path= you pass.
| Name | Summary |
|---|---|
mtb.data.fetch |
Download reference datasets into the data root, skipping those present. |
mtb.data.fetch_outputs |
Download precomputed run_all outputs for a tutorial dataset. |
mtb.data.fetchable |
Dataset ids mtb.data.fetch can download. |
mtb.data.fetch
¶
Download reference datasets into the data root, skipping those present.
| PARAMETERS | DESCRIPTION |
|---|---|
*datasets
|
Dataset ids to fetch;
type
|
data_path
|
Data root that holds the dataset folders;
type
|
quiet
|
Suppress the one "downloading ..." line per dataset.
type
|
| RETURNS | DESCRIPTION |
|---|---|
Path
|
The data root (not the dataset folder) - pass it as |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
A dataset id is not a release asset (the message lists the ids). |
RuntimeError
|
The downloaded archive lacks the |
OSError
|
The download failed; the message names the URL. |
Examples
>>> import multibench as mtb
>>> # the data root; D11 lands in <data_path>/D11/ (about 11 MB)
>>> root = mtb.data.fetch("D11")
>>> mtb.scan("D11", "vertical", data_path=root)
>>> mtb.data.fetch("D11", "D28", data_path="data", quiet=True)
Notes
What it covers. The tutorial reference sets published as release
assets (mtb.data.fetchable()); the progress line states each one's
approximate download size. The full collection is linked from the
scMultiBench README ("Get the data" in the installation guide).
Idempotence. A non-empty <data_path>/<dataset>/ counts as present
and is left alone, whatever it holds (even for an id that is not a
release asset); an empty leftover folder is removed and fetched again.
Atomic extraction. Each archive is unpacked into a temporary folder
under the data root and then moved into place. An archive entry that
would write outside that folder raises RuntimeError.
See Also
mtb.data.fetchable : the dataset ids this function can download.
mtb.data.fetch_outputs : precomputed run_all outputs for a tutorial dataset.
mtb.config.Config : data_path, the default data root.
mtb.data.fetch_outputs
¶
Download precomputed run_all outputs for a tutorial dataset.
Use it on a host without the method environments, such as Colab or a
laptop. evaluate and plot then run on the stored outputs.
| PARAMETERS | DESCRIPTION |
|---|---|
dataset
|
Dataset id with shipped outputs:
type
|
methods
|
Method ids that must be in the tree; checked only, the whole tree is downloaded either way.
type
|
data_path
|
Data root that holds the dataset folders;
type
|
quiet
|
Suppress the one "downloading ..." line.
type
|
| RETURNS | DESCRIPTION |
|---|---|
Path
|
|
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
|
KeyError
|
A name in |
RuntimeError
|
The archive lacks |
OSError
|
The download failed; the message names the URL. |
Examples
>>> import multibench as mtb
>>> out = mtb.data.fetch_outputs("D11")
>>> res = mtb.load_batch(out)
>>> res.summary[["method", "status", "ARI", "NMI"]]
Notes
Tree contents. Exactly what mtb.run_all writes and
mtb.load_batch reads: batch_result.json, long.csv,
summary.csv and one folder per method holding its embedding.h5.
Selecting methods. methods only checks the names; the whole tree
is downloaded. To restrict what is loaded, use
mtb.load_batch(out, methods=[...]).
Idempotence. When <data_path>/outputs/<dataset>/batch_result.json
exists nothing is downloaded. An empty leftover folder is replaced. A
non-empty one without batch_result.json raises RuntimeError.
Remove it and call again.
Atomic extraction. The archive is unpacked into a temporary folder
and then moved into place. The archive may be rooted at the tree itself
or at one folder (outputs-D11/, D11/); an entry that would write
outside the temporary folder raises RuntimeError.
See Also
mtb.load_batch : load the downloaded tree as a BatchResult.
mtb.run_all : produces the same tree from your own run.
mtb.data.fetch : download the input datasets themselves.
mtb.data.fetchable
¶
Dataset ids mtb.data.fetch can download.
The companion of mtb.available_datasets, which lists the ids with
stored results; the two sets overlap but are not the same.
| RETURNS | DESCRIPTION |
|---|---|
list of str
|
Dataset ids in natural order ( |
Examples
Notes
Overlap with stored results: D24 has published metric tables and
no downloadable file, D46 downloads and has no stored table. The
list is exactly the set of ids fetch() accepts.