Skip to content
Open In Colab

Quickstart

Open in Colab

Five steps take you from finding methods to plotting your run next to the stored results. Only step 3 runs a method. It needs Linux and a method environment. The other steps also work on macOS and Windows.

Words used on this site
embedding
A table with one row per cell and 10-50 numbers per cell, like adata.obsm["X_pca"].
category
How your cells were measured, for example vertical for CITE-seq from one experiment.
variant
One version of a method for one data type, such as RNA + ADT.
role
The name of one input file, such as rna or adt.
dataset folder
One folder that holds the input files of a dataset, such as rna.h5 and cty.csv.
raw counts
Integer counts per cell and feature, before any normalization.
demo dataset
A benchmark dataset that mtb.data.fetch downloads, such as D11 or D28.
stored tables
The metric tables that ship with the package: the benchmark's published scores and the package's own re-run scores for the demo datasets.
long table
A table with one row per method, dataset and metric.
cell order
The order of the rows in a method's output. For diagonal data most methods put the RNA cells first, then the ATAC cells. labels_for returns the label files in that order.

Summary

import multibench as mtb, pandas as pd

# 1. methods for unpaired RNA + ATAC (diagonal)
mtb.find_methods(category="diagonal", modalities=["rna", "atac"])

# 2. plot the stored table of dataset D28
df = mtb.load_results("diagonal", dataset="D28")
mtb.plot.bubble(df, metrics=["ARI", "NMI", "cLISI"], save="fig.pdf")

# 3. run a method on D28 (Linux, with the method's environment)
mtb.data.fetch("D28")
res = mtb.run("SCALEX", "diagonal",
              inputs=mtb.inputs_for("D28", "diagonal", "SCALEX"),
              out_dir="out/scalex_d28")

# 4. score it, with the Leiden backend of the stored tables
mtb.config.DEFAULT.leiden_flavor = "leidenalg"
labels = mtb.labels_for("D28", "diagonal", "SCALEX")
metrics = mtb.evaluate(res.output, labels=labels, metrics="all")

# 5. plot it next to the stored table
mine = mtb.to_long(metrics, method="SCALEX_rerun", dataset="D28",
                   category="diagonal")
mtb.plot.bubble(pd.concat([df, mine]), metrics=["ARI", "NMI", "cLISI"])

Step 1: Discover methods for your data

Pick the category that matches how your data was measured, then filter by modality:

  • vertical: several modalities measured in the same cells
  • diagonal: modalities measured in different cells
  • mosaic: several batches, only some sharing a modality
  • cross: several batches, each with every modality (RNA+ADT here)

For several samples or donors:

  • 10x Multiome samples: use vertical. Keep all cells in one folder and pass each cell's sample as batch= to run_all or evaluate.
  • CITE-seq donors: use vertical in the same way. The cross category also fits three donors, one file each: see the One file per batch tab.
import multibench as mtb

mtb.find_methods(category="diagonal", modalities=["rna", "atac"])
# -> ['SCALEX', 'Seurat_v5', 'scBridge', 'VIPCCA', 'Portal', 'uniPort',
#     'sciCAN', 'scJoint', 'GLUE', 'MultiMAP', 'iNMF', 'online_iNMF', 'Conos',
#     'Seurat_v3']

Each method reads ATAC either as peaks or as gene-activity scores. With a peak matrix only, keep the methods that read peaks:

mtb.find_methods(category="diagonal", atac="peak")
# -> ['Seurat_v5', 'GLUE', 'MultiMAP', 'Seurat_v3']

MultiMAP and Seurat_v3 also need gene activity. Seurat_v5 needs RNA and ATAC from the same cells. See My ATAC is peaks only.

Inspect a candidate, or rank candidates from the stored tables:

info = mtb.method_info("SCALEX")
info["env"], info["atac"], info["needs_labels"], info["reference"]["doi"]
# ('scmb_torch', 'gene_activity', False, '10.1038/s41467-022-33758-z')
info["runtime"]    # {..., 'tier': 'slow', 'worst_sec': 2233, ...}
mtb.recommend("diagonal", modalities=["rna", "atac"])
Details

find_methods also filters by task, needs_labels, runnable and tunable, and every filter must hold. The reference lists them.

recommend lists methods without stored scores last, with grand_score NaN, and names them in a warning. recommend(..., metrics="batch") ranks on the batch-correction metrics instead.

For CITE-seq donors, the vertical category has 14 RNA + ADT methods and the stored table D11. The cross category has 8 methods, which integrate the donors as batches, and the stored table D52. The cross methods read batches 1-3. UINMF reads only the first two.


Step 2: Plot the stored results

Load the stored table of a demo dataset and draw the scIB-style bubble table. load_results returns a long table: one row per method, dataset and metric.

df = mtb.load_results("diagonal", dataset="D28")          # 14 methods
fig = mtb.plot.bubble(df, metrics=["ARI", "NMI", "cLISI"], save="fig.pdf")
Details

source="published", the default, loads the benchmark's scIB tables. source="rerun" loads the package's own runs, which the category tutorials use. The two can hold different methods for one dataset: for D52, one published and eight re-run. mtb.results_coverage(category) lists what each holds.

load_results also takes methods, metrics and clustering (reference).

plot.bubble(..., aggregate="summary") ranks methods across several datasets by averaging their per-dataset ranks. Use it only for methods scored on every one of those datasets.


Step 3: Run a method

mtb.run runs one method in its own environment and loads its output. inputs_for returns the input files of a demo dataset. Install the method's environment first, on Linux: multibench env install --methods SCALEX --packed --run.

mtb.data.fetch("D28")  # 137 MB, downloaded once
res = mtb.run(
    method="SCALEX",
    category="diagonal",
    # {'rna': .../D28/rna.h5, 'atac_gas': .../D28/atac_gas.h5}
    inputs=mtb.inputs_for("D28", "diagonal", "SCALEX"),
    out_dir="out/scalex_d28",
)

res.output      # the joint embedding
res.cmd         # the command that ran
Details

inputs maps each role to a file or an AnnData. Inputs that are not in the package's .h5 layout are converted first, and convert=False skips this. A MuData raises an error: write it as a dataset folder with mtb.io.export_dataset first. Relative paths are fine.

cmd_template="{cmd}" runs the command in the current environment instead of the method's environment.

The RunResult also holds the method's captured stdout and stderr. A failed run raises RuntimeError with the end of that output.

dry_run=True returns the command without running it, on any platform. From the shell, --dry-run prints it. Drop the flag to run. Make the command on the machine that runs it, because it holds absolute paths.

terminal
multibench run --method SCALEX --category diagonal \
    --input rna=<data_path>/D28/rna.h5 \
    --input atac_gas=<data_path>/D28/atac_gas.h5 \
    --out-dir out/scalex_d28 --dry-run

multibench config prints the data path that <data_path> stands for.

Your own data

Give raw counts. On macOS and Windows you can prepare the folder, check it, preview the commands, score your own embedding and plot. Running a method needs Linux.

import anndata as ad
import multibench as mtb
import pandas as pd
import scanpy as sc

# adata: RNA counts in X, ADT counts in obsm["protein"], cell types in obs
adata = ad.read_h5ad("my_citeseq.h5ad")
mtb.io.export_dataset(adata, "data/MYCITE", rna="X", adt="obsm:protein",
                      labels="obs:celltype")
mtb.scan("MYCITE", "vertical", data_path="data", modalities=["rna", "adt"])
inputs = mtb.inputs_for("MYCITE", "vertical", "totalVI", data_path="data")
cmd = mtb.run("totalVI", "vertical", inputs=inputs, out_dir="out/totalVI",
              dry_run=True)
print(" ".join(cmd))  # the command, with this computer's paths

# compare two PCA embeddings: RNA and ADT
pc = adata.copy()
sc.pp.normalize_total(pc)
sc.pp.log1p(pc)
sc.pp.highly_variable_genes(pc, n_top_genes=2000)
sc.pp.pca(pc)
adata.obsm["X_pca"] = pc.obsm["X_pca"]
adt = sc.AnnData(adata.obsm["protein"])
sc.pp.normalize_total(adt)
sc.pp.log1p(adt)
sc.pp.pca(adt, n_comps=min(20, adt.n_vars - 1))
adata.obsm["X_adt"] = adt.obsm["X_pca"]
m_rna = mtb.evaluate(adata, labels="celltype", obsm="X_pca")
m_adt = mtb.evaluate(adata, labels="celltype", obsm="X_adt")
mine = pd.concat([
    mtb.to_long(m_rna, method="RNA PCA", dataset="MYCITE",
                category="vertical"),
    mtb.to_long(m_adt, method="ADT PCA", dataset="MYCITE",
                category="vertical"),
])
print(mine.pivot(index="method", columns="metric", values="value").round(2))
mtb.plot.bubble(mine)

With two rows, the darker one is the better of the two in each column. The printed table shows how far apart they are.

When the counts are in a layer, pass rna="layer:counts".

import anndata as ad
import multibench as mtb

adata = ad.read_h5ad("my_citeseq.h5ad")
mtb.io.export_dataset(adata, "data/MYCITE", rna="X", adt="obsm:protein",
                      labels="obs:celltype")
mtb.scan("MYCITE", "vertical", data_path="data", modalities=["rna", "adt"])
res = mtb.run_all("MYCITE", "vertical", data_path="data", out_dir="out/",
                  modalities=["rna", "adt"])
res.summary; res.plot()

run_all runs every method that scan marks runnable, scores each output and saves the results under out_dir.

10x Multiome, unpaired RNA + ATAC and one file per batch have their own recipes in Your own data as a dataset folder. mtb.describe_layout(category) prints the file names a category expects.


Step 4: Evaluate your run

mtb.evaluate scores an embedding with scIB metrics against the cell-type labels. Pass the dict of labels_for(dataset, category, method) as it is: its files are in the method's cell order. For diagonal data, the batch metrics iLISI, GC and ASW_batch show how well the RNA and ATAC cells mix.

# the Leiden backend of the stored tables
mtb.config.DEFAULT.leiden_flavor = "leidenalg"
metrics = mtb.evaluate(
    res.output,
    # {'rna_cty': ..., 'atac_cty': ...}
    labels=mtb.labels_for("D28", "diagonal", "SCALEX"),
    metrics="all",
)
# index: ARI, NMI, ASW, iASW, iF1, cLISI, ASW_batch, GC, iLISI   column: Value
Details

metrics=None, the default, computes every metric that applies except kBET. That is the clustering family, plus the batch family when there is a batch vector or several label files. See the metrics= vocabulary.

labels can also be a pandas.Series, an array, or an obs column name when the output is an AnnData. A dict in another order needs label_order=.

ARI, NMI and iF1 need a Leiden resolution sweep, which takes seconds to minutes. A metrics= list without those three skips it.

The stored tables were scored with leidenalg. The faster default, igraph, can move ARI by up to about 0.1.


Step 5: Plot your own result

to_long turns the scores into a long table. Plot it with the stored table of D28, under a method name that the table does not have yet.

The stored tables hold only the demo datasets. Rows from your own dataset go in a figure of their own.

import pandas as pd

mine = mtb.to_long(metrics, method="SCALEX_rerun", dataset="D28",
                   category="diagonal")

combined = pd.concat([df, mine], ignore_index=True)
mtb.plot.bubble(combined, metrics=["ARI", "NMI", "cLISI"], save="compare.pdf")
Details

A re-run differs slightly from the stored values. See Are my numbers comparable?

To rank your method against the stored rows, score it on the demo dataset of your category, such as D11 for CITE-seq or D28 for diagonal. Add your own method shows how. On your own dataset, plot methods scored on that dataset together: the run_all results on Linux, or your own PCA embeddings.

From the command line

Each step has a CLI equivalent. multibench <command> --help lists the flags.

terminal
multibench plot bubble --category diagonal --dataset D28 \
    --metrics ARI,NMI,cLISI --out fig.pdf
multibench evaluate --output out/scalex_d28/embedding.h5 \
    --method SCALEX --name SCALEX_rerun --dataset D28 --category diagonal \
    --leiden-flavor leidenalg --out mine.csv
multibench plot bubble --input mine.csv --category diagonal --dataset D28 \
    --metrics ARI,NMI,cLISI --out compare.pdf
Details
  • list, find, info, params, scan, layout, convert, fetch, config, run, run-all, evaluate, plot, cite and env mirror the Python API.
  • evaluate --method sets the label order. --name sets the row name. For your own method, give only --name and repeat --labels once per file, in your embedding's cell order.
  • scan prints a compact table. --columns all or --format csv|tsv|json gives every column.
  • run --dry-run and run-all --dry-run print the commands without running them.
  • --param METHOD:KEY=VALUE on run-all sets a hyperparameter. params METHOD lists them.
  • plot --input is repeatable. With --category, it adds your rows to the stored table.
  • Tables go to stdout. Progress and notes go to stderr.

Next steps