Skip to content

Compare

Read the stored metric tables and rank methods for a category. Score your own runs with mtb.config.DEFAULT.leiden_flavor = "leidenalg" before you compare them with these tables. Are my numbers comparable? explains why.

Name Summary
mtb.load_results Load the stored benchmark metric tables as one long table.
mtb.available_datasets Dataset ids that have stored results.
mtb.results_coverage Which (category, dataset, method) combinations have results, and where from.
mtb.recommend Rank methods from stored results, with the share of datasets each was scored on.
mtb.DegenerateRerunWarning A re-run row scored ARI ~0 where the published table scored well.

mtb.load_results

load_results(
    category: str | None = None,
    *,
    dataset: str | list[str] | None = None,
    methods: str | list[str] | None = None,
    metrics=None,
    clustering: str = "default",
    source: str = "published",
    result_path: Path | str | None = None,
) -> DataFrame

Load the stored benchmark metric tables as one long table.

Reads the published scIB tables, the package's re-run sweeps, or a long CSV of your own. Frames from any source concatenate with each other.

PARAMETERS DESCRIPTION
category

Integration category (vertical, diagonal, mosaic or cross); None = every category with tables for source.

type str default None

dataset

Dataset id(s), e.g. "D11" or ["D11", "D11s"]; None = every dataset of the category.

type str or list of str default None

methods

Method(s) to keep, alias tolerant and case-insensitive ("mofa+" -> MOFA2, "totalvi" -> totalVI); None = every method.

type str or list of str default None

metrics

None / "all" = every metric; or "clustering", "batch", or a list of codes such as ["ARI", "NMI"].

type None, str or list of str default None

clustering

Clustering variant of the published tables: metric.csv, metric_louvain.csv or metric_kmeans.csv.

type ('default', 'louvain', 'kmeans') default "default"

source

"published" (scIB tables), "rerun" (package sweeps) or "both". For a result_path file: a value of its source column; "published" / "both" keep every row.

type str default 'published'

result_path

A results root holding scib_metric/ and/or rerun/, or one long CSV file; None = the tables shipped in the package.

type path - like default None

RETURNS DESCRIPTION
DataFrame

One row per score, columns metric, value, method, dataset, category, clustering, source, plus scored_with when a user CSV has it. attrs["rerun_version"] holds the re-run stamp (Notes).

RAISES DESCRIPTION
FileNotFoundError

No table for a requested category, dataset or source, or no result_path.

KeyError

Unknown method name in methods; the message suggests a close match, if any.

ValueError

Unknown category, metrics, clustering or source, or a malformed result_path file.

TypeError

A positional argument after category, or a retired keyword (Notes).

Examples

>>> import multibench as mtb
>>> pub = mtb.load_results("diagonal", dataset="D28")   # published tables
>>> # package re-runs
>>> rr = mtb.load_results("diagonal", dataset="D28", source="rerun")
>>> both = mtb.load_results("cross", dataset="D52", source="both",
...                         metrics="batch")
>>> # your own rows
>>> mine = mtb.load_results(result_path="mine.csv", source="user")
Notes

Sources. source picks the tables under a results root (default mtb.config.DEFAULT.result_path, shipped in the package):

  • "published" - the scIB tables under result_path/scib_metric (shipped: vertical D11, diagonal D24/D25/D28, cross D52; none for mosaic);
  • "rerun" - the package's re-run sweeps result_path/rerun/long_all_<dataset>.csv (shipped: vertical D11/D11s, diagonal D28/D28s, mosaic D45/D45s, cross D52/D52s);
  • "both" - the concatenation; the source column tells the rows apart.

Re-run rows were scored by multibench 0.2.1's evaluate with the leidenalg backend and every cell type counted as isolated. Set mtb.config.DEFAULT.leiden_flavor = "leidenalg" to score comparable rows.

Your own file. result_path may be one long CSV with at least the columns metric, value, method, e.g. written by to_long(...).to_csv or BatchResult.save(). A file keeps whatever source / clustering values it carries; a missing column or a blank cell is filled with "user" / "default" per row. A scored_with column is kept, with NaN for the stored rows. source= then filters on the file's own source column:

  • "published" / "both" - keep every row;
  • "user" - the rows mtb.to_long wrote;
  • "rerun" - rerun and rerun-<version> rows (prefix match);
  • any value the file does not contain raises ValueError listing the ones present.

Method and metric names. A methods entry that is neither a method id nor a method in the loaded frame raises KeyError with a did-you-mean hint ("Unknown method Matlida. Did you mean Matilda?"). Method folders of the published tables are reported by method id (MOFA+ -> MOFA2, Seurat(WNN) -> Seurat_WNN); metric names read from the published tables or a file are canonical (iFI -> iF1).

Dataset ids. Every requested id must have a table: ["D11", "D99"] raises FileNotFoundError naming D99, the datasets that have tables and the result_path= hint. Ids are case-sensitive; with a category, an id that differs from a stored one only in case ("d52") is replaced by the stored spelling, with a UserWarning.

Metric selection. Uses the vocabulary of mtb.evaluate:

  • "clustering" - mtb.plot.CLUSTERING_METRICS (ARI, NMI, ASW, iASW, iF1, cLISI);
  • "batch" - mtb.plot.BATCH_METRICS (ASW_batch, GC, iLISI, kBET);
  • a list - exactly those codes (alias tolerant, ["ari"] -> ARI).

Every code must be one of mtb.catalog.known_metrics() or present in the frame: ["ZZZ"] raises ValueError listing both. An unknown token raises too ("dimension_reduction" is a list_tasks token, not a family).

Empty results. A known method or metric with no rows gives an empty frame and a UserWarning, not an error. Under source="published" the warning also says whether the re-run sweeps hold that method ("The re-run tables hold 1 dataset (D11). Pass source='rerun' or 'both'.").

Clustering variants. A result directory named <method>_louvain / <method>_kmeans is that variant: it is reported under the method's canonical id with clustering set to the suffix, and only when that variant is requested. clustering= reads the published tables only; re-run rows are "default".

Warnings about the tables. With "rerun" / "both", a re-run row whose ARI is ~0 while the published table scored the same method and dataset well emits a mtb.DegenerateRerunWarning (in the shipped sweeps: Conos on D28).

When the selection holds one method in the chosen source while the other source holds more, a UserWarning says so ("The published table for cross/D52 has one method, scMoMaT, ... Pass source="rerun" or "both"."). No warning when the other source has nothing more.

Re-run version. The sweep files stamp their rows rerun-<version>; the source column reads plain "rerun" and attrs["rerun_version"] keeps the version ("0.2.1"; a sorted tuple when files from several versions were loaded; None when no stamped row is present). Read attrs["rerun_version"] before pd.concat, which drops attrs that differ.

Retired keywords. method=, metric=, task= and family= still work, with a DeprecationWarning; metric_set= raises TypeError. The Changes page lists the replacements.

See Also

mtb.available_datasets : the dataset ids that have stored tables.

mtb.results_coverage : which methods each source covers, per dataset.

mtb.to_long : turns mtb.evaluate scores into this shape.

mtb.recommend : ranks methods from these tables.

mtb.plot.bubble : plots the frame.

mtb.available_datasets

available_datasets(
    category: str | None = None,
    *,
    source: str = "published",
    result_path: Path | str | None = None,
) -> list[str]

Dataset ids that have stored results.

mtb.data.fetchable lists the ids that can be downloaded.

PARAMETERS DESCRIPTION
category

Integration category; None = the union over all four.

type str default None

source

Which stored tables to look at (see mtb.load_results).

type ('published', 'rerun', 'both') default "published"

result_path

Results root; None = the tables shipped in the package.

type path - like default None

RETURNS DESCRIPTION
list of str

Sorted dataset ids.

RAISES DESCRIPTION
ValueError

Unknown category or source.

Examples

>>> import multibench as mtb
>>> mtb.available_datasets()  # ['D11', 'D24', 'D25', 'D28', 'D52']
>>> mtb.available_datasets("diagonal")  # ['D24', 'D25', 'D28']
>>> mtb.available_datasets(source="rerun")  # ['D11', 'D11s', 'D28', ...]
Notes

Downloadable datasets. An id listed here but not by mtb.data.fetchable() has metric tables and no data file to download.

What counts. Published ids are those holding at least one method's default-clustering table (metric.csv) - what load_results(category, dataset=...) loads. A category folder that does not exist (mosaic has no published tables) contributes nothing, without an error. A result_path that does not exist gives a UserWarning and [].

Retired keywords. metric_set= and clustering= are not parameters; passing either raises TypeError.

See Also

mtb.data.fetchable : the ids mtb.data.fetch can download.

mtb.results_coverage : which methods have results for each dataset.

mtb.load_results : loads the tables behind these ids.

mtb.results_coverage

results_coverage(
    category: str | None = None,
    *,
    source: str = "both",
    result_path: Path | str | None = None,
) -> DataFrame

Which (category, dataset, method) combinations have results, and where from.

Use it to choose the source to pass to mtb.load_results or mtb.recommend.

PARAMETERS DESCRIPTION
category

Integration category; None = all four.

type str default None

source

Which stored tables to scan; "both" = published and re-run.

type ('both', 'published', 'rerun') default "both"

result_path

Results root; None = the tables shipped in the package.

type path - like default None

RETURNS DESCRIPTION
DataFrame

One sorted row per distinct category, dataset, method, clustering, source; source is "published" or "rerun".

RAISES DESCRIPTION
ValueError

Unknown category or source.

Examples

>>> import multibench as mtb
>>> cov = mtb.results_coverage("cross")
>>> cov[cov.dataset == "D52"]     # scMoMaT (published) + 8 methods (rerun)
>>> cov.attrs["rerun_version"]    # '0.2.1'
>>> cov.groupby(["category", "source"]).method.nunique()
Notes

Clustering variants. The published tree is scanned for every clustering variant (default, louvain, kmeans), so a method that only exists as e.g. a _louvain directory shows up under clustering="louvain".

Empty categories. A category that has no tables raises nothing; it has no rows.

Warnings. The degenerate-row and one-method warnings of mtb.load_results are silenced here.

Re-run version. attrs["rerun_version"] holds the package version stamped on the re-run sweeps, as in mtb.load_results.

See Also

mtb.available_datasets : just the dataset ids.

mtb.load_results : the scores behind each row.

mtb.recommend

recommend(
    category: str,
    *,
    modalities: list[str] | None = None,
    atac: str | None = None,
    methods: list[str] | None = None,
    metrics=None,
    long_df: DataFrame | None = None,
    min_methods: int = 2,
    source: str = "published",
    result_path: Path | str | None = None,
) -> DataFrame

Rank methods from stored results, with the share of datasets each was scored on.

Scores each method on the stored metric tables (or long_df) and also lists every available method it could not score; the rules are in Notes.

PARAMETERS DESCRIPTION
category

Integration category to rank: vertical, diagonal, mosaic or cross.

type str

modalities

Keep only methods that consume all of these modalities, e.g. ["rna", "adt"]; tokens as in mtb.find_methods. None = no filter.

type list of str default None

atac

Keep only methods that read this ATAC representation: "peak" or "gene_activity"; None = no filter.

type str default None

methods

Rank only these methods (alias tolerant, case-insensitive), among themselves; None = every method.

type list of str default None

metrics

None = the "clustering" family (the benchmark's main ranking); or "batch", "all", or a list of codes.

type None, str or list of str default None

long_df

Frame to score (metric, value, method, dataset) instead of the stored tables, e.g. pd.concat([published, mine]).

type DataFrame default None

min_methods

Datasets with fewer methods than this are dropped.

type int default 2

source

Stored tables to load when long_df is not given; "both" averages the scores present in both.

type ('published', 'rerun', 'both') default "published"

result_path

Results root; None = the tables shipped in the package.

type path - like default None

RETURNS DESCRIPTION
DataFrame

One row per method, best first, unscored methods last. Read method, grand_score, coverage and datasets; all columns and attrs are in Notes.

RAISES DESCRIPTION
ValueError

Nothing left to rank (Notes lists the cases), or an unknown metrics value.

KeyError

Unknown method name in methods; the message suggests a close match, if any.

FileNotFoundError

No stored results for the category and source.

WARNS DESCRIPTION
UserWarning

Partial coverage, dropped or unscored methods, or igraph-scored rows; one message (Notes).

Examples

>>> import multibench as mtb
>>> r = mtb.recommend("vertical", modalities=["rna", "adt"])
>>> r[["method", "grand_score", "coverage", "datasets"]]
>>> r.attrs["not_scored"]                   # available, but no published rows
>>> mtb.recommend("diagonal", atac="peak")  # methods that read peaks
>>> mtb.recommend("cross", source="rerun")  # cross has one published method
Notes

Score. grand_score is the overall="mean_overall" score of mtb.plot.bar: within each dataset, the min-max scaled mean of the per-metric max-ranks, then averaged over the datasets the method was run on.

Ranking rules.

  • Only methods this package runs for the category are ranked - the set mtb.list_methods(category=...) lists. Other package methods in the table (MOFA2 or Multigrate in a cross table) are dropped before the within-dataset ranks are taken, and named in the warning and in attrs["dropped_methods"]. A method the package does not know (your own method in long_df) is kept.
  • A dataset holding fewer than min_methods methods is dropped. The min-max of a single method is always 1.0, so a lone method would win that dataset.
  • n_datasets / n_datasets_total / coverage say how much of the method x dataset matrix each score rests on.
  • Every method the package runs for the category (and modalities) that has no rows in the chosen source is still listed, after the scored rows, with grand_score NaN, n_datasets 0 and coverage 0.0. The re-run tables may hold more of these methods (source="rerun").

Columns.

  • method - the method id;
  • grand_score - the score above; NaN when unscored;
  • n_datasets / n_datasets_total - datasets the method was scored on / datasets kept for the ranking;
  • coverage - n_datasets / n_datasets_total;
  • needs_labels, runtime_tier, worst_sec, env, output_kind - package metadata for the category; None for ids that are not package methods (your own method, a result-dir token);
  • datasets - the dataset ids the score comes from, comma-joined; "" when unscored.

Which data the ranking describes. modalities and atac select methods; they do not change the datasets the scores come from. The shipped tables measured rna+adt (D11, D11s, D52, D52s) or rna+atac (D24, D25, D28, D28s, D45, D45s). When modalities or atac names a modality that none of the ranked datasets measured, a UserWarning says so.

The stored Seurat_v5 diagonal scores come from runs with a separate paired bridge dataset; mtb.scan marks Seurat_v5 runnable only when your RNA and ATAC files hold the same cells.

Frame attrs. frame.attrs records the choices the ranking was made under:

  • "metrics" - the family token, or the list of codes;
  • "family" - the token; None when a list was given;
  • "source" - "published" / "rerun" / "both", or "long_df";
  • "not_scored" (also under "missing") - the unscored method ids;
  • "dropped_methods" - package methods present in the table but not run by this package for the category.

Selections. metrics="batch" scores ASW_batch, GC, iLISI, kBET; "all" every metric present; a list exactly those codes (alias tolerant). methods is resolved as in mtb.load_results ("mofa+" -> MOFA2, "totalvi" -> totalVI), and the unscored line of the warning is restricted to the same set, so a requested method without rows is still reported as such. modalities and atac select methods with mtb.find_methods, under the same token rule.

Warning. One UserWarning with one line per finding summarises a requested modality the datasets did not measure, dropped methods and datasets, partial coverage and the unscored methods. The stored tables were clustered with leidenalg. A line also names the long_df methods clustered with igraph when ARI, NMI or iF1 ranks them against stored rows of their dataset.

Errors. ValueError is raised when:

  • no dataset holds min_methods methods (when the other stored source holds more methods for the category the message says so - cross/D52: pass source='rerun');
  • the frame has none of the requested metrics (the message lists the metrics it does have);
  • every row belongs to a method the package does not run for the category;
  • methods=, modalities= or atac= leaves no scored method;
  • an unknown modality token or atac value, or a representation token that contradicts atac (the rule of mtb.find_methods);
  • a metrics token or code is unknown.

Retired keywords. task= / family= still work as metrics=<token>, with a DeprecationWarning.

See Also

mtb.load_results : the stored tables it ranks.

mtb.results_coverage : which methods each source covers.

mtb.plot.bar : plots the same overall score per method.

mtb.DegenerateRerunWarning

Bases: UserWarning

A re-run row scored ARI ~0 where the published table scored well.

Emitted by mtb.load_results and mtb.recommend with source="rerun" or "both".

Notes

When it fires:

  • a re-run row scored ARI below 0.01 where the published table scored the same category, dataset and method above 0.2. Such a row most likely comes from a failed re-run, for example a collapsed embedding or a wrong label order, not from the method itself;
  • never for a result_path file or a long_df frame passed to mtb.recommend;
  • in the shipped sweeps: Conos on D28.

What to do: drop the row before ranking (the message names the filter, here df[df.method != 'Conos']); once you have decided how to treat those rows, silence it with warnings.simplefilter("ignore", mtb.DegenerateRerunWarning).