Skip to content

Load and plot results

mtb.load_results reads the stored metric tables into one long table. mtb.plot.bubble draws the scIB-style bubble table from it.

Summary

import multibench as mtb

df = mtb.load_results("vertical", dataset="D11")
fig = mtb.plot.bubble(df, metrics=["ARI", "NMI", "ASW", "cLISI"],
                      title="Vertical integration, D11", save="d11.pdf")

The same panel from the shell:

terminal
multibench plot bubble --category vertical --dataset D11 \
    --metrics ARI,NMI,ASW,cLISI --out d11.pdf

Step 1: Load a results table

load_results returns a long table: one row per method, dataset and metric. Every plotting function takes this table.

import multibench as mtb

df = mtb.load_results(
    "vertical",            # vertical | diagonal | mosaic | cross
    dataset="D11",         # a list for several; omit for all
    source="published",    # published | rerun | both
)
df.head()
#   metric     value      method dataset  category clustering     source
# 0    NMI  0.792936  Multigrate     D11  vertical    default  published
# 1    ARI  0.759446  Multigrate     D11  vertical    default  published
# 2    ASW  0.656011  Multigrate     D11  vertical    default  published
# ...
Details

methods= and metrics= narrow the table and accept aliases, such as "mofa+" for MOFA2. metrics= takes the same values as in evaluate.

result_path= reads a long CSV you saved, or another results root, instead of the shipped tables.

Published and re-run tables

source= picks the table. "published", the default, holds the benchmark's scIB tables, "rerun" the package's own re-run sweeps, and "both" joins the two. mtb.available_datasets(category, source=...) lists the datasets each holds.

Details

The datasets in each source:

Category published rerun
vertical D11 D11, D11s
diagonal D24, D25, D28 D28, D28s
mosaic none D45, D45s
cross D52 D52, D52s

The two sources can hold different methods for one dataset. For D52, published holds one method and rerun eight. mtb.results_coverage(category) lists what each holds.

The re-run table was scored with this package's evaluate and leidenalg. The published table follows the benchmark's own scoring. A method in both gets different numbers from each: see Are my numbers comparable?.

A re-run row with an ARI near 0, where the published table scored the same method well, triggers mtb.DegenerateRerunWarning. The warning names the row and the line that drops it.

Clustering variants

clustering= selects the clustering behind the clustering metrics: "default" (Leiden), "louvain" or "kmeans". In the published tables, diagonal has all three variants for D24, D25 and D28. vertical and cross have default only.

Details

The re-run tables are default only. With source="rerun", every row is tagged default whatever clustering= asks for.


Step 2: Draw the bubble table

mtb.plot.bubble draws one row per method, best on top, and one column per metric. It returns a Matplotlib Figure.

fig = mtb.plot.bubble(
    df,
    metrics=["ARI", "NMI", "ASW", "cLISI"],   # column order within each family
    title="Vertical integration, D11",
    save="d11_bubble.pdf",
)
Details

The reference lists every option. These are used most often:

  • methods=[...] keeps only these methods, before ranking.
  • order=[...] puts these methods first, in this order. The rest follow, best first.
  • show_language=False hides the language chips and the L badge.
  • na="skip" silences the warning about missing metrics.

Reading the figure

Circle size shows the rank within a column. A bigger circle is better. Fill colour is scaled within each column of this figure. The lightest fill is the lowest value in that column. Each metric family starts with an Overall bar, and rows are ordered by the mean of those bars.

Details

Blue columns are the dimension-reduction and clustering metrics, and green columns the batch-correction metrics.

A grey column holds the same value in every row. In a summary panel it holds the same mean rank. A figure with one method is grey in every column. A warning and the footnote name the grey columns.

A dash marks a metric not computed for that method. Its Overall averages the metrics it has, and a warning names each such method.

The chips give each method's language, Py or R. ? marks a name the package does not know, such as your own method or a renamed re-run. L marks a method that uses cell-type labels. Its scores are not comparable with the rows without labels.

Individual vs. summary panels

aggregate="dataset", the default, plots one dataset's values. aggregate="summary" ranks methods within each dataset, then averages the ranks across datasets. A summary compares methods scored on the same datasets. require_complete=True keeps only the methods scored on every dataset.

df = mtb.load_results("vertical", dataset=["D11", "D11s"], source="rerun")
fig = mtb.plot.bubble(df, metrics=["ARI", "NMI", "ASW", "cLISI"],
                      aggregate="summary", require_complete=True,
                      title="Vertical integration, D11 + D11s")
Details

D11s, D28s, D45s and D52s are random subsamples of the full datasets, scored in the re-run tables. They cannot be rebuilt. For a summary with your own method, use the full datasets.

Under "dataset", several datasets are averaged per method. Each row label shows the method's dataset, or how many datasets it averages (Matilda ยท 2 ds). A warning says so. When the datasets share no method, plot each dataset on its own. If they hold the same cells, set one dataset name first: df.assign(dataset="MYDATA").

Under "summary" with the default overall="rank", a method missing from a dataset scores rank 0 there, which pulls it down. A warning lists these methods.

overall="rank", the bubble default, and "mean_overall", the bar default, combine datasets differently and can order methods differently. Pass the same value to plot.bubble and plot.bar for the same order.

Audit the numbers before drawing

mtb.plot.build_table takes the same selection arguments and returns the numbers behind the figure as a BubbleTable.

table = mtb.plot.build_table(df, metrics=["ARI", "NMI", "ASW", "cLISI"],
                             aggregate="summary")
table.methods   # row order, best first
table.ranks     # per-column ranks, highest = best

Summary bars

mtb.plot.bar draws one bar per method with its Overall score across datasets, for one metric family.

fig = mtb.plot.bar(df, group="clustering", save="bar.pdf")
Details

With several datasets and the default overall="mean_overall", one dot per dataset sits behind each bar.


Add your own method

Score your run with mtb.evaluate, reshape it with mtb.to_long, and concatenate it with a stored table. Score with the leidenalg backend, as Are my numbers comparable? explains.

d11_panel.py
import pandas as pd
import multibench as mtb

# the Leiden backend of the stored tables
mtb.config.DEFAULT.leiden_flavor = "leidenalg"

bench = mtb.load_results("vertical", dataset="D11", source="rerun")
# emb: your embedding, cells x dims, in the order of the D11 labels
mine = mtb.to_long(mtb.evaluate(emb, labels=mtb.labels_for("D11")["cty"]),
                   method="MyMethod", dataset="D11", category="vertical")

both = pd.concat([bench, mine])
metrics = ["ARI", "NMI", "ASW", "cLISI"]
fig = mtb.plot.bubble(both, metrics=metrics,
                      title="Vertical integration, D11 + MyMethod",
                      save="d11_vertical.pdf")
table = mtb.plot.build_table(both, metrics=metrics)
print(table.methods.index("MyMethod") + 1)   # its row in the figure
Details

Give your run a name the stored table does not use. Duplicate (method, dataset, metric) rows raise instead of being averaged.

Add mine["needs_labels"] = True to show the L badge for a method that uses cell-type labels.

Try it with a PCA of D11

To try the recipe without a method of your own, use a PCA of D11's RNA as emb. Put these lines above bench = ... in d11_panel.py, then run the script.

import scanpy as sc

mtb.data.fetch("D11")
d = mtb.config.DEFAULT.data_path / "D11"
rna = mtb.io.read_canonical(d / "rna.h5")
sc.pp.normalize_total(rna)
sc.pp.log1p(rna)
sc.pp.highly_variable_genes(rna, n_top_genes=2000)
sc.pp.pca(rna)
emb = rna.obsm["X_pca"]

From the shell

The same figure without Python, with your rows saved by mine.to_csv("mine.csv", index=False):

terminal
multibench plot bubble \
    --category vertical --dataset D11 --source rerun \
    --input mine.csv \
    --metrics ARI,NMI,ASW,cLISI \
    --out d11_vertical.pdf
Details

--aggregate, --overall, --require-complete, --methods and --title match the Python keywords. --result-path DIR reads another results root.

--input takes a long CSV or a run_all output directory and can be repeated. Alone it plots your table. With --category, your rows are added to the stored table.

multibench plot bar --group clustering draws the summary bars. The --out extension sets the format (.pdf, .png, ...).