Compare¶
Read the stored metric tables and rank methods for a category. Score your own
runs with mtb.config.DEFAULT.leiden_flavor = "leidenalg" before you compare
them with these tables.
Are my numbers comparable?
explains why.
| Name | Summary |
|---|---|
mtb.load_results |
Load the stored benchmark metric tables as one long table. |
mtb.available_datasets |
Dataset ids that have stored results. |
mtb.results_coverage |
Which (category, dataset, method) combinations have results, and where from. |
mtb.recommend |
Rank methods from stored results, with the share of datasets each was scored on. |
mtb.DegenerateRerunWarning |
A re-run row scored ARI ~0 where the published table scored well. |
mtb.load_results
¶
load_results(
category: str | None = None,
*,
dataset: str | list[str] | None = None,
methods: str | list[str] | None = None,
metrics=None,
clustering: str = "default",
source: str = "published",
result_path: Path | str | None = None,
) -> DataFrame
Load the stored benchmark metric tables as one long table.
Reads the published scIB tables, the package's re-run sweeps, or a long CSV of your own. Frames from any source concatenate with each other.
| PARAMETERS | DESCRIPTION |
|---|---|
category
|
Integration category (
type
|
dataset
|
Dataset id(s), e.g.
type
|
methods
|
Method(s) to keep, alias tolerant and case-insensitive (
type
|
metrics
|
type
|
clustering
|
Clustering variant of the published tables:
type
|
source
|
type
|
result_path
|
A results root holding
type
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
One row per score, columns |
| RAISES | DESCRIPTION |
|---|---|
FileNotFoundError
|
No table for a requested category, dataset or source, or no |
KeyError
|
Unknown method name in |
ValueError
|
Unknown |
TypeError
|
A positional argument after |
Examples
>>> import multibench as mtb
>>> pub = mtb.load_results("diagonal", dataset="D28") # published tables
>>> # package re-runs
>>> rr = mtb.load_results("diagonal", dataset="D28", source="rerun")
>>> both = mtb.load_results("cross", dataset="D52", source="both",
... metrics="batch")
>>> # your own rows
>>> mine = mtb.load_results(result_path="mine.csv", source="user")
Notes
Sources. source picks the tables under a results root
(default mtb.config.DEFAULT.result_path, shipped in the package):
"published"- the scIB tables underresult_path/scib_metric(shipped: vertical D11, diagonal D24/D25/D28, cross D52; none for mosaic);"rerun"- the package's re-run sweepsresult_path/rerun/long_all_<dataset>.csv(shipped: vertical D11/D11s, diagonal D28/D28s, mosaic D45/D45s, cross D52/D52s);"both"- the concatenation; thesourcecolumn tells the rows apart.
Re-run rows were scored by multibench 0.2.1's evaluate with the
leidenalg backend and every cell type counted as isolated. Set
mtb.config.DEFAULT.leiden_flavor = "leidenalg" to score comparable
rows.
Your own file. result_path may be one long CSV with at least the
columns metric, value, method, e.g. written by to_long(...).to_csv
or BatchResult.save(). A file keeps whatever source /
clustering values it carries; a missing column or a blank cell is
filled with "user" / "default" per row. A scored_with column
is kept, with NaN for the stored rows. source= then filters on the
file's own source column:
"published"/"both"- keep every row;"user"- the rowsmtb.to_longwrote;"rerun"-rerunandrerun-<version>rows (prefix match);- any value the file does not contain raises
ValueErrorlisting the ones present.
Method and metric names. A methods entry that is neither a
method id nor a method in the loaded frame raises KeyError with a
did-you-mean hint ("Unknown method Matlida. Did you mean
Matilda?"). Method folders of the published tables are reported by
method id (MOFA+ -> MOFA2, Seurat(WNN) -> Seurat_WNN);
metric names read from the published tables or a file are canonical
(iFI -> iF1).
Dataset ids. Every requested id must have a table: ["D11",
"D99"] raises FileNotFoundError naming D99, the datasets that
have tables and the result_path= hint. Ids are case-sensitive; with
a category, an id that differs from a stored one only in case
("d52") is replaced by the stored spelling, with a UserWarning.
Metric selection. Uses the vocabulary of mtb.evaluate:
"clustering"-mtb.plot.CLUSTERING_METRICS(ARI, NMI, ASW, iASW, iF1, cLISI);"batch"-mtb.plot.BATCH_METRICS(ASW_batch, GC, iLISI, kBET);- a list - exactly those codes (alias tolerant,
["ari"]-> ARI).
Every code must be one of mtb.catalog.known_metrics() or present in
the frame: ["ZZZ"] raises ValueError listing both. An unknown
token raises too ("dimension_reduction" is a list_tasks token,
not a family).
Empty results. A known method or metric with no rows gives an empty
frame and a UserWarning, not an error. Under source="published"
the warning also says whether the re-run sweeps hold that method
("The re-run tables hold 1 dataset (D11). Pass source='rerun' or
'both'.").
Clustering variants. A result directory named <method>_louvain /
<method>_kmeans is that variant: it is reported under the method's
canonical id with clustering set to the suffix, and only when that
variant is requested. clustering= reads the published tables only;
re-run rows are "default".
Warnings about the tables. With "rerun" / "both", a re-run
row whose ARI is ~0 while the published table scored the same method
and dataset well emits a mtb.DegenerateRerunWarning (in the shipped
sweeps: Conos on D28).
When the selection holds one method in the chosen source while the
other source holds more, a UserWarning says so ("The published
table for cross/D52 has one method, scMoMaT, ... Pass source="rerun" or
"both"."). No warning when the other source has nothing more.
Re-run version. The sweep files stamp their rows
rerun-<version>; the source column reads plain "rerun" and
attrs["rerun_version"] keeps the version ("0.2.1"; a sorted
tuple when files from several versions were loaded; None when no
stamped row is present). Read attrs["rerun_version"] before
pd.concat, which drops attrs that differ.
Retired keywords. method=, metric=, task= and
family= still work, with a DeprecationWarning; metric_set=
raises TypeError. The Changes page lists the replacements.
See Also
mtb.available_datasets : the dataset ids that have stored tables.
mtb.results_coverage : which methods each source covers, per dataset.
mtb.to_long : turns mtb.evaluate scores into this shape.
mtb.recommend : ranks methods from these tables.
mtb.plot.bubble : plots the frame.
mtb.available_datasets
¶
available_datasets(
category: str | None = None,
*,
source: str = "published",
result_path: Path | str | None = None,
) -> list[str]
Dataset ids that have stored results.
mtb.data.fetchable lists the ids that can be downloaded.
| PARAMETERS | DESCRIPTION |
|---|---|
category
|
Integration category;
type
|
source
|
Which stored tables to look at (see
type
|
result_path
|
Results root;
type
|
| RETURNS | DESCRIPTION |
|---|---|
list of str
|
Sorted dataset ids. |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
Unknown |
Examples
>>> import multibench as mtb
>>> mtb.available_datasets() # ['D11', 'D24', 'D25', 'D28', 'D52']
>>> mtb.available_datasets("diagonal") # ['D24', 'D25', 'D28']
>>> mtb.available_datasets(source="rerun") # ['D11', 'D11s', 'D28', ...]
Notes
Downloadable datasets. An id listed here but not by
mtb.data.fetchable() has metric tables and no data file to download.
What counts. Published ids are those holding at least one method's
default-clustering table (metric.csv) - what
load_results(category, dataset=...) loads. A category folder that
does not exist (mosaic has no published tables) contributes
nothing, without an error. A result_path that does not exist gives
a UserWarning and [].
Retired keywords. metric_set= and clustering= are not
parameters; passing either raises TypeError.
See Also
mtb.data.fetchable : the ids mtb.data.fetch can download.
mtb.results_coverage : which methods have results for each dataset.
mtb.load_results : loads the tables behind these ids.
mtb.results_coverage
¶
results_coverage(
category: str | None = None,
*,
source: str = "both",
result_path: Path | str | None = None,
) -> DataFrame
Which (category, dataset, method) combinations have results, and where from.
Use it to choose the source to pass to mtb.load_results or
mtb.recommend.
| PARAMETERS | DESCRIPTION |
|---|---|
category
|
Integration category;
type
|
source
|
Which stored tables to scan;
type
|
result_path
|
Results root;
type
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
One sorted row per distinct |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
Unknown |
Examples
>>> import multibench as mtb
>>> cov = mtb.results_coverage("cross")
>>> cov[cov.dataset == "D52"] # scMoMaT (published) + 8 methods (rerun)
>>> cov.attrs["rerun_version"] # '0.2.1'
>>> cov.groupby(["category", "source"]).method.nunique()
Notes
Clustering variants. The published tree is scanned for every
clustering variant (default, louvain, kmeans), so a method that only
exists as e.g. a _louvain directory shows up under
clustering="louvain".
Empty categories. A category that has no tables raises nothing; it has no rows.
Warnings. The degenerate-row and one-method warnings of
mtb.load_results are silenced here.
Re-run version. attrs["rerun_version"] holds the package version
stamped on the re-run sweeps, as in mtb.load_results.
See Also
mtb.available_datasets : just the dataset ids.
mtb.load_results : the scores behind each row.
mtb.recommend
¶
recommend(
category: str,
*,
modalities: list[str] | None = None,
atac: str | None = None,
methods: list[str] | None = None,
metrics=None,
long_df: DataFrame | None = None,
min_methods: int = 2,
source: str = "published",
result_path: Path | str | None = None,
) -> DataFrame
Rank methods from stored results, with the share of datasets each was scored on.
Scores each method on the stored metric tables (or long_df) and also
lists every available method it could not score; the rules are in Notes.
| PARAMETERS | DESCRIPTION |
|---|---|
category
|
Integration category to rank:
type
|
modalities
|
Keep only methods that consume all of these modalities, e.g.
type
|
atac
|
Keep only methods that read this ATAC representation:
type
|
methods
|
Rank only these methods (alias tolerant, case-insensitive), among
themselves;
type
|
metrics
|
type
|
long_df
|
Frame to score (
type
|
min_methods
|
Datasets with fewer methods than this are dropped.
type
|
source
|
Stored tables to load when
type
|
result_path
|
Results root;
type
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
One row per method, best first, unscored methods last. Read
|
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
Nothing left to rank (Notes lists the cases), or an unknown
|
KeyError
|
Unknown method name in |
FileNotFoundError
|
No stored results for the category and source. |
| WARNS | DESCRIPTION |
|---|---|
UserWarning
|
Partial coverage, dropped or unscored methods, or igraph-scored rows; one message (Notes). |
Examples
>>> import multibench as mtb
>>> r = mtb.recommend("vertical", modalities=["rna", "adt"])
>>> r[["method", "grand_score", "coverage", "datasets"]]
>>> r.attrs["not_scored"] # available, but no published rows
>>> mtb.recommend("diagonal", atac="peak") # methods that read peaks
>>> mtb.recommend("cross", source="rerun") # cross has one published method
Notes
Score. grand_score is the overall="mean_overall" score of
mtb.plot.bar: within each dataset, the min-max scaled mean of the
per-metric max-ranks, then averaged over the datasets the method was run
on.
Ranking rules.
- Only methods this package runs for the category are ranked - the set
mtb.list_methods(category=...)lists. Other package methods in the table (MOFA2 or Multigrate in a cross table) are dropped before the within-dataset ranks are taken, and named in the warning and inattrs["dropped_methods"]. A method the package does not know (your own method inlong_df) is kept. - A dataset holding fewer than
min_methodsmethods is dropped. The min-max of a single method is always 1.0, so a lone method would win that dataset. n_datasets/n_datasets_total/coveragesay how much of the method x dataset matrix each score rests on.- Every method the package runs for the category (and
modalities) that has no rows in the chosen source is still listed, after the scored rows, withgrand_scoreNaN,n_datasets0 andcoverage0.0. The re-run tables may hold more of these methods (source="rerun").
Columns.
method- the method id;grand_score- the score above; NaN when unscored;n_datasets/n_datasets_total- datasets the method was scored on / datasets kept for the ranking;coverage-n_datasets / n_datasets_total;needs_labels,runtime_tier,worst_sec,env,output_kind- package metadata for the category;Nonefor ids that are not package methods (your own method, a result-dir token);datasets- the dataset ids the score comes from, comma-joined;""when unscored.
Which data the ranking describes. modalities and atac select
methods; they do not change the datasets the scores come from. The
shipped tables measured rna+adt (D11, D11s, D52, D52s) or rna+atac (D24,
D25, D28, D28s, D45, D45s). When modalities or atac names a
modality that none of the ranked datasets measured, a UserWarning
says so.
The stored Seurat_v5 diagonal scores come from runs with a separate paired
bridge dataset; mtb.scan marks Seurat_v5 runnable only when your RNA and
ATAC files hold the same cells.
Frame attrs. frame.attrs records the choices the ranking was
made under:
"metrics"- the family token, or the list of codes;"family"- the token;Nonewhen a list was given;"source"-"published"/"rerun"/"both", or"long_df";"not_scored"(also under"missing") - the unscored method ids;"dropped_methods"- package methods present in the table but not run by this package for the category.
Selections. metrics="batch" scores ASW_batch, GC, iLISI, kBET;
"all" every metric present; a list exactly those codes (alias
tolerant). methods is resolved as in mtb.load_results
("mofa+" -> MOFA2, "totalvi" -> totalVI), and the unscored line
of the warning is restricted to the same set, so a requested method
without rows is still reported as such. modalities and atac
select methods with mtb.find_methods, under the same token rule.
Warning. One UserWarning with one line per finding summarises
a requested modality the datasets did not measure, dropped methods and
datasets, partial coverage and the unscored methods. The stored tables
were clustered with leidenalg. A line also names the long_df methods
clustered with igraph when ARI, NMI or iF1 ranks them against stored
rows of their dataset.
Errors. ValueError is raised when:
- no dataset holds
min_methodsmethods (when the other stored source holds more methods for the category the message says so - cross/D52:pass source='rerun'); - the frame has none of the requested metrics (the message lists the metrics it does have);
- every row belongs to a method the package does not run for the category;
methods=,modalities=oratac=leaves no scored method;- an unknown modality token or
atacvalue, or a representation token that contradictsatac(the rule ofmtb.find_methods); - a
metricstoken or code is unknown.
Retired keywords. task= / family= still work as
metrics=<token>, with a DeprecationWarning.
See Also
mtb.load_results : the stored tables it ranks.
mtb.results_coverage : which methods each source covers.
mtb.plot.bar : plots the same overall score per method.
mtb.DegenerateRerunWarning
¶
Bases: UserWarning
A re-run row scored ARI ~0 where the published table scored well.
Emitted by mtb.load_results and mtb.recommend with
source="rerun" or "both".
Notes
When it fires:
- a re-run row scored ARI below 0.01 where the published table scored the same category, dataset and method above 0.2. Such a row most likely comes from a failed re-run, for example a collapsed embedding or a wrong label order, not from the method itself;
- never for a
result_pathfile or along_dfframe passed tomtb.recommend; - in the shipped sweeps: Conos on D28.
What to do: drop the row before ranking (the message names the
filter, here df[df.method != 'Conos']); once you have decided how to
treat those rows, silence it with
warnings.simplefilter("ignore", mtb.DegenerateRerunWarning).