Load and plot results¶
mtb.load_results
reads the stored metric tables into one long table.
mtb.plot.bubble draws
the scIB-style bubble table from it.
Summary
import multibench as mtb
df = mtb.load_results("vertical", dataset="D11")
fig = mtb.plot.bubble(df, metrics=["ARI", "NMI", "ASW", "cLISI"],
title="Vertical integration, D11", save="d11.pdf")
The same panel from the shell:
Step 1: Load a results table¶
load_results returns a long table: one row per method, dataset and
metric. Every plotting function takes this table.
import multibench as mtb
df = mtb.load_results(
"vertical", # vertical | diagonal | mosaic | cross
dataset="D11", # a list for several; omit for all
source="published", # published | rerun | both
)
df.head()
# metric value method dataset category clustering source
# 0 NMI 0.792936 Multigrate D11 vertical default published
# 1 ARI 0.759446 Multigrate D11 vertical default published
# 2 ASW 0.656011 Multigrate D11 vertical default published
# ...
Details
methods= and metrics= narrow the table and accept aliases, such as
"mofa+" for MOFA2. metrics= takes the same values as in
evaluate.
result_path= reads a long CSV you saved, or another results root,
instead of the shipped tables.
Published and re-run tables¶
source= picks the table. "published", the default, holds the
benchmark's scIB tables, "rerun" the package's own re-run sweeps, and
"both" joins the two. mtb.available_datasets(category, source=...)
lists the datasets each holds.
Details
The datasets in each source:
| Category | published |
rerun |
|---|---|---|
| vertical | D11 | D11, D11s |
| diagonal | D24, D25, D28 | D28, D28s |
| mosaic | none | D45, D45s |
| cross | D52 | D52, D52s |
The two sources can hold different methods for one dataset. For D52,
published holds one method and rerun eight.
mtb.results_coverage(category) lists what each holds.
The re-run table was scored with this package's evaluate and
leidenalg. The published table follows the benchmark's own scoring. A
method in both gets different numbers from each: see
Are my numbers comparable?.
A re-run row with an ARI near 0, where the published table scored the
same method well, triggers mtb.DegenerateRerunWarning. The warning
names the row and the line that drops it.
Clustering variants¶
clustering= selects the clustering behind the clustering metrics:
"default" (Leiden), "louvain" or "kmeans". In the published tables,
diagonal has all three variants for D24, D25 and D28. vertical
and cross have default only.
Details
The re-run tables are default only. With source="rerun", every row
is tagged default whatever clustering= asks for.
Step 2: Draw the bubble table¶
mtb.plot.bubble draws one row per method, best on top, and one column per
metric. It returns a Matplotlib Figure.
fig = mtb.plot.bubble(
df,
metrics=["ARI", "NMI", "ASW", "cLISI"], # column order within each family
title="Vertical integration, D11",
save="d11_bubble.pdf",
)
Details
The reference lists every option. These are used most often:
methods=[...]keeps only these methods, before ranking.order=[...]puts these methods first, in this order. The rest follow, best first.show_language=Falsehides the language chips and theLbadge.na="skip"silences the warning about missing metrics.
Reading the figure¶
Circle size shows the rank within a column. A bigger circle is better. Fill colour is scaled within each column of this figure. The lightest fill is the lowest value in that column. Each metric family starts with an Overall bar, and rows are ordered by the mean of those bars.
Details
Blue columns are the dimension-reduction and clustering metrics, and green columns the batch-correction metrics.
A grey column holds the same value in every row. In a summary panel it holds the same mean rank. A figure with one method is grey in every column. A warning and the footnote name the grey columns.
A dash marks a metric not computed for that method. Its Overall averages the metrics it has, and a warning names each such method.
The chips give each method's language, Py or R. ? marks a name
the package does not know, such as your own method or a renamed
re-run. L marks a method that uses cell-type labels. Its scores are
not comparable with the rows without labels.
Individual vs. summary panels¶
aggregate="dataset", the default, plots one dataset's values.
aggregate="summary" ranks methods within each dataset, then averages the
ranks across datasets. A summary compares methods scored on the same
datasets. require_complete=True keeps only the methods scored on every
dataset.
df = mtb.load_results("vertical", dataset=["D11", "D11s"], source="rerun")
fig = mtb.plot.bubble(df, metrics=["ARI", "NMI", "ASW", "cLISI"],
aggregate="summary", require_complete=True,
title="Vertical integration, D11 + D11s")
Details
D11s, D28s, D45s and D52s are random subsamples of the full datasets, scored in the re-run tables. They cannot be rebuilt. For a summary with your own method, use the full datasets.
Under "dataset", several datasets are averaged per method. Each row
label shows the method's dataset, or how many datasets it averages
(Matilda ยท 2 ds). A warning says so. When the datasets share no
method, plot each dataset on its own. If they hold the same cells, set
one dataset name first: df.assign(dataset="MYDATA").
Under "summary" with the default overall="rank", a method missing
from a dataset scores rank 0 there, which pulls it down. A warning
lists these methods.
overall="rank", the bubble default, and "mean_overall", the bar
default, combine datasets differently and can order methods
differently. Pass the same value to plot.bubble and plot.bar for
the same order.
Audit the numbers before drawing¶
mtb.plot.build_table takes the same selection arguments and returns the
numbers behind the figure as a
BubbleTable.
table = mtb.plot.build_table(df, metrics=["ARI", "NMI", "ASW", "cLISI"],
aggregate="summary")
table.methods # row order, best first
table.ranks # per-column ranks, highest = best
Summary bars¶
mtb.plot.bar draws one bar per method with its Overall score across
datasets, for one metric family.
Details
With several datasets and the default overall="mean_overall", one
dot per dataset sits behind each bar.
Add your own method¶
Score your run with mtb.evaluate, reshape it with mtb.to_long, and
concatenate it with a stored table. Score with the leidenalg backend, as
Are my numbers comparable?
explains.
import pandas as pd
import multibench as mtb
# the Leiden backend of the stored tables
mtb.config.DEFAULT.leiden_flavor = "leidenalg"
bench = mtb.load_results("vertical", dataset="D11", source="rerun")
# emb: your embedding, cells x dims, in the order of the D11 labels
mine = mtb.to_long(mtb.evaluate(emb, labels=mtb.labels_for("D11")["cty"]),
method="MyMethod", dataset="D11", category="vertical")
both = pd.concat([bench, mine])
metrics = ["ARI", "NMI", "ASW", "cLISI"]
fig = mtb.plot.bubble(both, metrics=metrics,
title="Vertical integration, D11 + MyMethod",
save="d11_vertical.pdf")
table = mtb.plot.build_table(both, metrics=metrics)
print(table.methods.index("MyMethod") + 1) # its row in the figure
Details
Give your run a name the stored table does not use. Duplicate (method, dataset, metric) rows raise instead of being averaged.
Add mine["needs_labels"] = True to show the L badge for a method
that uses cell-type labels.
Try it with a PCA of D11¶
To try the recipe without a method of your own, use a PCA of D11's RNA
as emb. Put these lines above bench = ... in d11_panel.py, then run
the script.
import scanpy as sc
mtb.data.fetch("D11")
d = mtb.config.DEFAULT.data_path / "D11"
rna = mtb.io.read_canonical(d / "rna.h5")
sc.pp.normalize_total(rna)
sc.pp.log1p(rna)
sc.pp.highly_variable_genes(rna, n_top_genes=2000)
sc.pp.pca(rna)
emb = rna.obsm["X_pca"]
From the shell¶
The same figure without Python, with your rows saved by
mine.to_csv("mine.csv", index=False):
multibench plot bubble \
--category vertical --dataset D11 --source rerun \
--input mine.csv \
--metrics ARI,NMI,ASW,cLISI \
--out d11_vertical.pdf
Details
--aggregate, --overall, --require-complete, --methods and
--title match the Python keywords. --result-path DIR reads another
results root.
--input takes a long CSV or a run_all output directory and can be
repeated. Alone it plots your table. With --category, your rows are
added to the stored table.
multibench plot bar --group clustering draws the summary bars. The
--out extension sets the format (.pdf, .png, ...).