Skip to content

Catalog (mtb.catalog)

The benchmark's methods, datasets and metrics tables as DataFrames. canonical_id and canonical_metric turn any known spelling of a method id or metric code into the canonical one.

Name Summary
mtb.catalog.methods Table of the benchmark's methods, one row per method.
mtb.catalog.datasets Table of the benchmark's datasets, joined with the stored results.
mtb.catalog.metrics Table describing the scIB metrics: one row per metric code.
mtb.catalog.canonical_id Return the canonical method id for any known spelling.
mtb.catalog.canonical_metric Return the canonical code for a metric name ("ari" -> "ARI").
mtb.catalog.known_metrics The canonical metric codes the package knows, in family order.

mtb.catalog.methods

methods(files_dir: Path | str | None = None) -> DataFrame

Table of the benchmark's methods, one row per method.

PARAMETERS DESCRIPTION
files_dir

Folder holding method.csv; None = the package's shipped files/.

type Path | str | None default None

RETURNS DESCRIPTION
DataFrame

One row per method. Read method, needs_labels, atac and categories; all columns are listed in Notes.

Examples

>>> import multibench as mtb
>>> m = mtb.catalog.methods()
>>> m[m.needs_labels][["method", "categories"]]
>>> m[m.categories.map(lambda c: "vertical" in c)].method.tolist()
Notes

Column reference.

  • method - the name as spelled in method.csv;
  • canonical_id - the method id (mtb.catalog.canonical_id);
  • language - 'python' or 'r' (lower-cased);
  • deep_learning - 'Yes' / 'No', as in the CSV;
  • atac - 'peak', 'gene_activity' or None;
  • output - 'embedding' or 'graph';
  • needs_labels - bool, the method needs cell-type labels;
  • categories / tasks - lists of integration categories and tasks.

Package values. For every method id, needs_labels, atac, categories and tasks come from the package's method definitions, which mtb.method_info and mtb.scan read. A row without a method id keeps the CSV values. language, deep_learning and output come from method.csv.

See Also

mtb.list_methods : the method ids, optionally per category. mtb.method_info : everything the package knows about one method. mtb.catalog.canonical_id : the normalisation behind the canonical_id column.

mtb.catalog.datasets

datasets(
    files_dir: Path | str | None = None, *, category: str | None = None
) -> DataFrame

Table of the benchmark's datasets, joined with the stored results.

PARAMETERS DESCRIPTION
files_dir

Folder holding dataset.csv; None = the package's shipped files/.

type Path | str | None default None

category

Keep only datasets with stored results in this integration category (vertical, diagonal, mosaic or cross); None = all.

type str | None default None

RETURNS DESCRIPTION
DataFrame

One row per dataset id. Read dataset, category and has_results; all columns are listed in Notes.

RAISES DESCRIPTION
ValueError

Unknown category; the message lists the valid ones.

Examples

>>> import multibench as mtb
>>> ds = mtb.catalog.datasets()
>>> ds[ds.has_results][["dataset", "category"]]
>>> mtb.catalog.datasets(category="vertical").dataset.tolist()
Notes

Column reference.

  • dataset - the id (D11, SD15, ...);
  • dataset name - a duplicate of dataset, kept for one release for older callers;
  • simulated - bool, ids starting with SD;
  • category - the integration categories whose stored results (published or re-run) contain the dataset, ";"-joined; None when no stored results exist;
  • has_results - bool, a stored metric table (mtb.load_results) covers it.

A dataset.csv passed through files_dir may also fill the descriptive columns assay, tissue, n_cells, n_batches and source; each one that holds a value is added. The shipped file leaves them empty.

Row set. The rows are the union of dataset.csv and every id with stored results (mtb.available_datasets(source="both")), so D11s/D28s/D45s/D52s and D24 (published tables only) are listed although dataset.csv does not name them. Ids missing from the CSV are appended after it, in natural order.

Subsamples. D11s, D28s, D45s and D52s are random subsamples of the full datasets, used for the second re-run sweep; they cannot be fetched. D28s holds 60% of D28's cells.

When the columns are computed. category and has_results are derived at each call from mtb.available_datasets; a missing or unreadable result tree leaves them empty instead of raising.

See Also

mtb.available_datasets : dataset ids with stored results, per category. mtb.data.fetchable : dataset ids that can be downloaded. mtb.load_results : the stored metric tables themselves.

mtb.catalog.metrics

metrics(files_dir: Path | str | None = None) -> DataFrame

Table describing the scIB metrics: one row per metric code.

PARAMETERS DESCRIPTION
files_dir

Folder holding metric_full.csv; None = the package's shipped files/.

type Path | str | None default None

RETURNS DESCRIPTION
DataFrame

One row per scIB metric, with the CSV's columns (metric and description in the shipped file).

Examples

>>> import multibench as mtb
>>> tab = mtb.catalog.metrics()
>>> tab.set_index("metric").loc["kBET", "description"]
Notes

Rows. The ten scIB metrics ARI, NMI, ASW, iASW, iF1, cLISI, ASW_batch, GC, iLISI, kBET; header whitespace is stripped. The canonical code vocabulary, including PCR from the published tables, is mtb.catalog.known_metrics.

See Also

mtb.catalog.known_metrics : the canonical metric codes. mtb.evaluate : computes these metrics for an embedding.

mtb.catalog.canonical_id

canonical_id(name: str, *, strict: bool = False) -> str

Return the canonical method id for any known spelling.

PARAMETERS DESCRIPTION
name

Any spelling of a method name ("MOFA+", "totalvi", "Seurat(WNN)").

type str

strict

True = raise for a name that is neither an alias nor a method id; False = return the folded name.

type bool default False

RETURNS DESCRIPTION
str

The canonical id (an alias target or a method id); for an unknown name, the input with spaces and dots collapsed to _.

RAISES DESCRIPTION
KeyError

strict=True and the name is neither an alias nor a method id.

Examples

>>> import multibench as mtb
>>> mtb.catalog.canonical_id("MOFA+"), mtb.catalog.canonical_id("totalvi")
('MOFA2', 'totalVI')
>>> mtb.catalog.canonical_id("my method")          # unknown: separators folded
'my_method'
Notes

Resolution order. The first rule that matches wins:

  1. the alias table, case-insensitive ("MOFA+" -> "MOFA2", "Seurat(WNN)" -> "Seurat_WNN");
  2. a case-folded match against the method ids, after collapsing spaces and dots to _ ("totalvi" -> "totalVI", "scmomat" -> "scMoMaT");
  3. for a name the package does not know (a result-directory token, a user's own method name): the input with separators collapsed to _, unchanged in case.

Strict mode. The error is the message mtb.method_info and mtb.scan give, e.g. "Unknown method Matlida. Did you mean Matilda? mtb.list_methods() shows all methods.". An alias-table hit is returned without the method-id check, even with strict=True: "Seurat v4" -> "Seurat_v4", which is not a method id. The default is lenient: result directories and user frames can hold names the package does not know.

See Also

mtb.list_methods : the method ids. mtb.catalog.canonical_metric : the same normalisation for metric codes.

mtb.catalog.canonical_metric

canonical_metric(code: str, *, strict: bool = False) -> str | None

Return the canonical code for a metric name ("ari" -> "ARI").

PARAMETERS DESCRIPTION
code

A metric name in any spelling the package or scIB uses ("kbet", "isolated_label_f1").

type str

strict

True = raise for a code not in known_metrics(); False = return an unknown code unchanged (stripped).

type bool default False

RETURNS DESCRIPTION
str or None

The canonical code; None for None, an empty string or the string "nan" (a blank cell, which callers drop).

RAISES DESCRIPTION
ValueError

strict=True and the code is not a known metric (the message lists them).

Examples

>>> import multibench as mtb
>>> mtb.catalog.canonical_metric("kbet")
'kBET'
>>> mtb.catalog.canonical_metric("isolated_label_f1")
'iF1'
>>> mtb.catalog.canonical_metric("my_score")        # unknown: kept as given
'my_score'
Notes

Spellings. Matching ignores case and surrounding whitespace: "iFI", "if1" and "isolated_label_f1" all give "iF1". The raw scIB long names "ARI_cluster/label", "NMI_cluster/label", "ASW_label" and "isolated_label_silhouette" map to ARI, NMI, ASW and iASW.

Unknown codes. The default returns an unknown code stripped but otherwise unchanged, so a user frame can carry a metric the package does not know. With strict=True the error reads "unknown metric 'nope'; valid: ['ARI', 'NMI', ...]".

See Also

mtb.catalog.known_metrics : the codes strict=True accepts. mtb.catalog.canonical_id : the same normalisation for method names.

mtb.catalog.known_metrics

known_metrics() -> list[str]

The canonical metric codes the package knows, in family order.

RETURNS DESCRIPTION
list of str

The clustering / bio-conservation codes (mtb.plot.CLUSTERING_METRICS), the batch-correction codes (mtb.plot.BATCH_METRICS), then PCR (principal-component regression, in some published tables).

Examples

>>> import multibench as mtb
>>> mtb.catalog.known_metrics()
['ARI', 'NMI', 'ASW', 'iASW', 'iF1', 'cLISI', 'ASW_batch', 'GC', 'iLISI',
 'kBET', 'PCR']
See Also

mtb.catalog.canonical_metric : maps any spelling to these codes; strict=True accepts only them. mtb.catalog.metrics : the description of each scIB metric.