Catalog (mtb.catalog)¶
The benchmark's methods, datasets and metrics tables as DataFrames.
canonical_id and canonical_metric turn any known spelling of a method id
or metric code into the canonical one.
| Name | Summary |
|---|---|
mtb.catalog.methods |
Table of the benchmark's methods, one row per method. |
mtb.catalog.datasets |
Table of the benchmark's datasets, joined with the stored results. |
mtb.catalog.metrics |
Table describing the scIB metrics: one row per metric code. |
mtb.catalog.canonical_id |
Return the canonical method id for any known spelling. |
mtb.catalog.canonical_metric |
Return the canonical code for a metric name ("ari" -> "ARI"). |
mtb.catalog.known_metrics |
The canonical metric codes the package knows, in family order. |
mtb.catalog.methods
¶
Table of the benchmark's methods, one row per method.
| PARAMETERS | DESCRIPTION |
|---|---|
files_dir
|
Folder holding
type
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
One row per method. Read |
Examples
>>> import multibench as mtb
>>> m = mtb.catalog.methods()
>>> m[m.needs_labels][["method", "categories"]]
>>> m[m.categories.map(lambda c: "vertical" in c)].method.tolist()
Notes
Column reference.
method- the name as spelled inmethod.csv;canonical_id- the method id (mtb.catalog.canonical_id);language-'python'or'r'(lower-cased);deep_learning-'Yes'/'No', as in the CSV;atac-'peak','gene_activity'orNone;output-'embedding'or'graph';needs_labels- bool, the method needs cell-type labels;categories/tasks- lists of integration categories and tasks.
Package values. For every method id, needs_labels,
atac, categories and tasks come from the package's method
definitions, which mtb.method_info and mtb.scan read. A row
without a method id keeps the CSV values. language, deep_learning and
output come from method.csv.
See Also
mtb.list_methods : the method ids, optionally per category.
mtb.method_info : everything the package knows about one method.
mtb.catalog.canonical_id : the normalisation behind the canonical_id column.
mtb.catalog.datasets
¶
Table of the benchmark's datasets, joined with the stored results.
| PARAMETERS | DESCRIPTION |
|---|---|
files_dir
|
Folder holding
type
|
category
|
Keep only datasets with stored results in this integration category
(
type
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
One row per dataset id. Read |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
Unknown |
Examples
>>> import multibench as mtb
>>> ds = mtb.catalog.datasets()
>>> ds[ds.has_results][["dataset", "category"]]
>>> mtb.catalog.datasets(category="vertical").dataset.tolist()
Notes
Column reference.
dataset- the id (D11,SD15, ...);dataset name- a duplicate ofdataset, kept for one release for older callers;simulated- bool, ids starting withSD;category- the integration categories whose stored results (published or re-run) contain the dataset,";"-joined;Nonewhen no stored results exist;has_results- bool, a stored metric table (mtb.load_results) covers it.
A dataset.csv passed through files_dir may also fill the
descriptive columns assay, tissue, n_cells, n_batches and
source; each one that holds a value is added. The shipped file
leaves them empty.
Row set. The rows are the union of dataset.csv and every id with
stored results (mtb.available_datasets(source="both")), so
D11s/D28s/D45s/D52s and D24 (published tables only)
are listed although dataset.csv does not name them. Ids missing from
the CSV are appended after it, in natural order.
Subsamples. D11s, D28s, D45s and D52s are random subsamples of the full datasets, used for the second re-run sweep; they cannot be fetched. D28s holds 60% of D28's cells.
When the columns are computed. category and has_results are
derived at each call from mtb.available_datasets; a missing or
unreadable result tree leaves them empty instead of raising.
See Also
mtb.available_datasets : dataset ids with stored results, per category. mtb.data.fetchable : dataset ids that can be downloaded. mtb.load_results : the stored metric tables themselves.
mtb.catalog.metrics
¶
Table describing the scIB metrics: one row per metric code.
| PARAMETERS | DESCRIPTION |
|---|---|
files_dir
|
Folder holding
type
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
One row per scIB metric, with the CSV's columns ( |
Examples
>>> import multibench as mtb
>>> tab = mtb.catalog.metrics()
>>> tab.set_index("metric").loc["kBET", "description"]
Notes
Rows. The ten scIB metrics ARI, NMI, ASW, iASW, iF1, cLISI,
ASW_batch, GC, iLISI, kBET; header whitespace is stripped. The
canonical code vocabulary, including PCR from the published tables,
is mtb.catalog.known_metrics.
See Also
mtb.catalog.known_metrics : the canonical metric codes. mtb.evaluate : computes these metrics for an embedding.
mtb.catalog.canonical_id
¶
Return the canonical method id for any known spelling.
| PARAMETERS | DESCRIPTION |
|---|---|
name
|
Any spelling of a method name (
type
|
strict
|
type
|
| RETURNS | DESCRIPTION |
|---|---|
str
|
The canonical id (an alias target or a method id); for an unknown
name, the input with spaces and dots collapsed to |
| RAISES | DESCRIPTION |
|---|---|
KeyError
|
|
Examples
>>> import multibench as mtb
>>> mtb.catalog.canonical_id("MOFA+"), mtb.catalog.canonical_id("totalvi")
('MOFA2', 'totalVI')
>>> mtb.catalog.canonical_id("my method") # unknown: separators folded
'my_method'
Notes
Resolution order. The first rule that matches wins:
- the alias table, case-insensitive (
"MOFA+"->"MOFA2","Seurat(WNN)"->"Seurat_WNN"); - a case-folded match against the method ids, after collapsing spaces
and dots to
_("totalvi"->"totalVI","scmomat"->"scMoMaT"); - for a name the package does not know (a result-directory token, a
user's own method name): the input with separators collapsed to
_, unchanged in case.
Strict mode. The error is the message mtb.method_info and
mtb.scan give, e.g. "Unknown method Matlida. Did you mean Matilda?
mtb.list_methods() shows all methods.". An
alias-table hit is returned without the method-id check, even with
strict=True: "Seurat v4" -> "Seurat_v4", which is not a
method id. The default is lenient: result directories and user
frames can hold names the package does not know.
See Also
mtb.list_methods : the method ids. mtb.catalog.canonical_metric : the same normalisation for metric codes.
mtb.catalog.canonical_metric
¶
Return the canonical code for a metric name ("ari" -> "ARI").
| PARAMETERS | DESCRIPTION |
|---|---|
code
|
A metric name in any spelling the package or scIB uses
(
type
|
strict
|
type
|
| RETURNS | DESCRIPTION |
|---|---|
str or None
|
The canonical code; |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
|
Examples
>>> import multibench as mtb
>>> mtb.catalog.canonical_metric("kbet")
'kBET'
>>> mtb.catalog.canonical_metric("isolated_label_f1")
'iF1'
>>> mtb.catalog.canonical_metric("my_score") # unknown: kept as given
'my_score'
Notes
Spellings. Matching ignores case and surrounding whitespace:
"iFI", "if1" and "isolated_label_f1" all give "iF1". The
raw scIB long names "ARI_cluster/label", "NMI_cluster/label",
"ASW_label" and "isolated_label_silhouette" map to ARI,
NMI, ASW and iASW.
Unknown codes. The default returns an unknown code stripped but
otherwise unchanged, so a user frame can carry a metric the package
does not know. With strict=True the error reads "unknown metric
'nope'; valid: ['ARI', 'NMI', ...]".
See Also
mtb.catalog.known_metrics : the codes strict=True accepts.
mtb.catalog.canonical_id : the same normalisation for method names.
mtb.catalog.known_metrics
¶
The canonical metric codes the package knows, in family order.
| RETURNS | DESCRIPTION |
|---|---|
list of str
|
The clustering / bio-conservation codes ( |
Examples
>>> import multibench as mtb
>>> mtb.catalog.known_metrics()
['ARI', 'NMI', 'ASW', 'iASW', 'iF1', 'cLISI', 'ASW_batch', 'GC', 'iLISI',
'kBET', 'PCR']
See Also
mtb.catalog.canonical_metric : maps any spelling to these codes; strict=True accepts only them.
mtb.catalog.metrics : the description of each scIB metric.