API reference¶
The package is imported as import multibench as mtb. The tables below list
every public function by step. Each name links to its reference entry, with
the signature, parameters and examples.
The five steps¶
A benchmark run has five steps: discover a method, prepare the data, run the method, score its output and compare the scores. Only step 3 needs Linux. The other steps also work on macOS and Windows.
import multibench as mtb
# score with the Leiden backend of the stored tables
mtb.config.DEFAULT.leiden_flavor = "leidenalg"
# Step 1: discover
mtb.find_methods(category="vertical", modalities=["rna", "adt"])
# Step 2: prepare (raw counts)
mtb.io.export_dataset(mdata, "data/MYCITE", rna="rna", adt="adt",
labels="rna:celltype")
# Step 3: run and score every method (Linux) ...
batch = mtb.run_all("MYCITE", "vertical", out_dir="runs/MYCITE",
data_path="data")
# ... or run one method
inputs = mtb.inputs_for("MYCITE", "vertical", "Matilda", data_path="data")
res = mtb.run("Matilda", "vertical", inputs=inputs, out_dir="runs/Matilda")
# Step 4: score
labels = mtb.labels_for("MYCITE", "vertical", "Matilda", data_path="data")
scores = mtb.evaluate(res.output, labels=labels)
# Step 5: compare
stored = mtb.load_results("vertical", dataset="D11")
mtb.plot.bubble(stored)
Details
Are my numbers comparable?
explains the leiden_flavor line and its shell flag,
--leiden-flavor leidenalg.
To place your method in a stored table, score it on the same demo dataset, such as D11 for CITE-seq. Add your own method shows the code.
mtb.to_long
turns the scores into a long table, the shape
mtb.load_results
returns. A long table holds one score per row, with its method, dataset
and metric. Concatenate the two tables and draw them with
mtb.plot.bubble.
Step 1: Discover¶
Find the methods that fit your data, what each one needs and how to cite it. These calls work offline. A method has one variant for each category and set of modalities it supports.
| entry point | what it does |
|---|---|
mtb.list_methods |
Method ids, optionally for one category. |
mtb.find_methods |
Method ids that match every filter you give: category, task, modalities, ATAC representation and more. |
mtb.method_info |
What the package knows about one method: environment, variants, runtime, GPU needs and reference. |
mtb.params_for |
The hyperparameters a method variant accepts, with their defaults. |
mtb.describe_layout |
How to lay out your own dataset folder for a category. |
mtb.cite |
Citation text or BibTeX for the benchmark and the methods you ran. |
mtb.list_tasks |
The valid task values. |
mtb.list_categories |
The valid category values, each with a short description. |
Step 2: Prepare data¶
Put your matrices into a dataset folder and find the files a method reads.
Give raw counts: the methods normalise the data themselves. A role names one
input of a method, such as rna, adt or cty.
mtb.run takes (method, category). mtb.inputs_for and mtb.labels_for
take (dataset, category, method). mtb.scan takes (dataset, category).
In mtb.io, mod names a MuData modality and modality names the role.
| entry point | what it does |
|---|---|
mtb.io.export_dataset |
Write an AnnData, a MuData or loose matrices as a dataset folder. |
mtb.io.to_canonical |
Convert one matrix to the canonical .h5 file, features x cells. |
mtb.io.read_canonical |
Read a canonical .h5 back as an AnnData, cells x features. |
mtb.io.normalize_peak_names |
Write a copy of a canonical .h5 with ATAC peak names spelled the way the methods expect. |
mtb.inputs_for |
{role: path} of the files a method variant reads from a dataset folder. |
mtb.labels_for |
{name: path} of a dataset's cell-type label files. With category and method, only the files that method reads, in the method's cell order. |
mtb.data.fetch |
Download the named reference datasets unless they are already present. |
mtb.data.fetch_outputs |
Download precomputed run_all outputs for a tutorial dataset. |
mtb.data.fetchable |
The dataset ids fetch can download. |
Step 3: Run¶
Run methods in their own environments, on Linux. mtb.scan shows what can
run before you start.
| entry point | what it does |
|---|---|
mtb.scan |
Which methods can run on a dataset, why the others cannot, and the command each would run. |
mtb.run |
Run one method and load its output, a RunResult. |
mtb.run_all |
Run and score every applicable method. Returns a BatchResult. |
mtb.sweep |
Run one method once per value of one hyperparameter, each in its own out_dir. |
mtb.load_batch |
Reload a saved run_all result. |
mtb.BatchResult |
What run_all returns: a summary table, a long table, a figure and rescoring. |
mtb.env.status |
Which method environments are installed on this machine. |
mtb.env.plan |
Which environments a set of methods needs. multibench env plan also shows their sizes. |
mtb.env.install |
Install those environments. A dry run until dry_run=False. |
mtb.env.doctor |
Which needed environments are installed and which are missing. multibench env doctor also prints the install command. |
mtb.env.recipe |
The build recipe of one method's environment. |
mtb.env.env_prefix |
The folder of an installed environment, the one mtb.run enters, or None. |
Step 4: Score¶
Compute scIB metrics for a run output.
| entry point | what it does |
|---|---|
mtb.evaluate |
scIB metrics of a run output, scored against cell-type labels and, when given, batch labels. |
mtb.to_long |
Reshape evaluate's result into the long results table. |
Step 5: Compare¶
Load the stored metric tables, rank methods and draw figures.
| entry point | what it does |
|---|---|
mtb.load_results |
The stored metric tables as one long table. |
mtb.available_datasets |
Dataset ids with stored results. |
mtb.results_coverage |
Which (category, dataset, method) combinations have results, and from which source. |
mtb.recommend |
Rank methods for a category from the stored results, with the share of datasets each method was scored on. |
mtb.DegenerateRerunWarning |
Warning class for a re-run row that most likely failed without an error. |
mtb.plot.bubble |
Bubble table of a long results table. |
mtb.plot.bar |
Each method's overall score across datasets, as bars. |
mtb.plot.build_table |
The numbers behind a bubble table, without drawing it. |
mtb.plot.BubbleTable |
What build_table returns: per-family matrices, ranks and row order. |
mtb.catalog.methods |
The methods table as a DataFrame. |
mtb.catalog.datasets |
The datasets table, including which datasets have stored results. |
mtb.catalog.metrics |
The metrics table, with a description of each metric. |
mtb.catalog.canonical_id |
The canonical method id for any known spelling. |
mtb.catalog.canonical_metric |
The canonical metric code for any known spelling, such as "ari" -> "ARI". |
mtb.catalog.known_metrics |
The metric codes, in family order. |
Configuration and errors¶
The config page lists the environment variables.
| entry point | what it does |
|---|---|
mtb.config.Config |
Paths and settings such as data_path, envs_dir and leiden_flavor. Every function reads mtb.config.DEFAULT. |
mtb.AmbiguousVariantError |
Raised when a method has several variants and neither modalities= nor the dataset's files select one. |
The same story from the shell¶
Each multibench subcommand wraps one Python function and takes the same
argument names. multibench <command> --help lists the flags. run and
run-all need Linux.
# how to lay out your data
multibench layout vertical
# write your raw counts as a dataset folder, then check it
multibench convert my.h5ad data/MYCITE --rna X --adt obsm:protein \
--labels obs:celltype
multibench scan MYCITE --category vertical --data-path data
# the hyperparameters --param accepts
multibench params Matilda
# run one method and score it, or run and score every method
multibench run --method Matilda --category vertical --out-dir runs/Matilda \
--input rna=data/MYCITE/rna.h5 --input adt=data/MYCITE/adt.h5 \
--input cty=data/MYCITE/cty.csv
multibench evaluate --output runs/Matilda/embedding.h5 \
--labels data/MYCITE/cty.csv --leiden-flavor leidenalg \
--method Matilda --dataset MYCITE --category vertical \
--out runs/Matilda/long.csv
multibench run-all MYCITE --category vertical --data-path data \
--out-dir runs/MYCITE --leiden-flavor leidenalg
# your runs, then a stored table
multibench plot bubble --input runs/MYCITE --out mine.pdf
multibench plot bubble --category vertical --dataset D11 --out d11.pdf
multibench cite Matilda MOFA2
Details
- Tables, ids and commands go to stdout. Progress, warnings and errors go
to stderr, so
2>/dev/nullhides them. scanandrun-all --dry-runclip long text to 80 characters.--columns allor--format csv|tsv|jsonprints every column in full, includingcommand.run --dry-runprints the command without running it.run-all --dry-runneeds no--out-dir. Its commands then show<out_dir>.scan --strictexits with1when no requested row is runnable. With--methods, it exits with1when any named method has no runnable row. It also exits with1while the method scripts are not fetched.--allow-atac-mismatchonscanandrun-allcounts a method as runnable when its ATAC file holds the other representation or peak names the method cannot read.--assume-gpuonscanandrun-all --dry-runchecks a GPU-node job from a login node.params METHODlists the hyperparameters per variant. Set one with--param KEY=VALUEonrun, or--param METHOD:KEY=VALUEonrun-all.evaluate --metricstakes a family (clustering,batch,all) or a comma-separated list of codes. Repeat--labelsfor several label files: they are stacked in the order given, one batch per file.--batch CSVonevaluateandrun-allgives one batch id per cell.plot --inputtogether with--categoryadds your rows to the stored table.--datasetfilters both, so your rows must be for a dataset the stored table has.- Pass
--categorytoconvert: the ATAC file name depends on the category. - The
envcommands are dry runs until--run. configlists the paths in use and where each one comes from.info METHODshows what one method needs.fetchdownloads demo datasets, their stored outputs (--outputs) or the method scripts (--scripts).- The exit code is
0on success,1on a runtime error,2on a usage error, and3whenrun-allfinished but a method failed, or a method named in--methodswas skipped. A line on stderr names those methods. A skipped method that--methodsdid not name does not set3, and thereasoncolumn ofsummary.csvsays why it did not run.MULTIBENCH_DEBUG=1prints the traceback of an error.
Errors and warnings¶
An invalid argument raises a standard exception whose message names the valid values or the fix.
Details
KeyError: unknown method id, with a did-you-mean hint.ValueError: unknowncategory,task,atacor modality token. The message lists the valid values.TypeError: a single string where a list is expected, such asmethods="StabMap". The message shows the list form.FileNotFoundError: missing dataset folder or file. The message lists the folders that exist, or the path and the working directory.FileExistsError:mtb.io.export_datasetormultibench convertwould replace a file. Passoverwrite=True, or--overwriteon the command line.mtb.AmbiguousVariantError: a method with several variants that neithermodalities=nor the dataset's files select. It is aValueErrorand also aKeyError.OSError: the method's environment is not installed, the method needs an NVIDIA GPU and this computer has none, or a download ofmtb.data.fetchorfetch_outputsfailed. On Linux, a missing environment gets the install command. On macOS and Windows, the message says that methods run only on Linux.
The package warns with UserWarning, and with DeprecationWarning for
old spellings.
mtb.DegenerateRerunWarning,
a UserWarning subclass, marks a re-run row that most likely failed
without an error. You can filter it by class.
The metrics= vocabulary¶
evaluate,
load_results,
recommend and
BatchResult.rescore take the same selector: "clustering", "batch",
"all", or a list of metric codes such as ["ARI", "NMI"].
Details
The three family names select these codes:
"clustering": ARI, NMI, ASW, iASW, iF1, cLISI (mtb.plot.CLUSTERING_METRICS)."batch": ASW_batch, GC, iLISI, kBET (mtb.plot.BATCH_METRICS)."all": both families.evaluatethen needs batch labels.
With the default None, evaluate computes the clustering family, plus
ASW_batch, GC and iLISI when batch labels are available. load_results
keeps every metric, and recommend ranks on the clustering family.
evaluate computes kBET only when a list names it.
A single code still goes in a list: metrics=["ARI"]. A task name or an
unknown code raises ValueError listing the valid values.
Method environments¶
Each method runs in a linux-64 environment, installed with mtb.env.install.
Compatible methods share one environment. mtb.method_info(m)["env"] names
it. On macOS and Windows, methods do not run: run and install work only
as dry runs.
Details
The methods' own scripts run unmodified. The package builds the command, converts the inputs and loads the output.
mtb.run downloads the method scripts on first use. On a machine without
network, run multibench fetch --scripts first. Add --ref, or set
MULTIBENCH_SCRIPTS_REF, to fetch a given commit or tag.
multibench config prints the commit in use, and every run records it
as scripts_commit.
On macOS and Windows, mtb.env.install(..., dry_run=False) raises
RuntimeError unless force=True (--force on the command line).
Run modes: prefix and conda¶
mtb.run enters each method's environment directly or through conda run.
How the environment is entered
explains both modes and MULTIBENCH_RUN_MODE=conda|prefix. A cmd_template
with {cmd} replaces this, and {env_cmd} keeps it.
Archive flavours: cpu and gpu¶
Five torch environments have a CUDA build and a smaller CPU-only build.
mtb.env.install picks the CPU build when
mtb.env.host_has_gpu
finds no NVIDIA GPU.
Details
The five environments are env_sciPENN, matilda, scmb_scjoint,
scmb_torch and scmb_scmm2. Every other environment has a single
build, the same archive for CPU and GPU hosts.
flavor="cpu" or flavor="gpu" overrides the default "auto"
(--flavor on the command line). Both builds install under the
environment's own name. mtb.env.status() reports the installed build as
flavor: cpu, gpu, or single for an environment with one build.
GPU or CPU?¶
Six methods need an NVIDIA GPU. mtb.method_info(m)["requires_gpu"] says
which. Without a GPU, mtb.run raises OSError before starting, and
mtb.scan gives the reason. The other methods run without a GPU.
Details
The six methods are scBridge, moETM, SMILE, sciCAN, UnitedNet and iPOLNG.
method_info(m)["gpu"] says how a method uses a GPU: required,
used when present, not used or unknown. multibench info METHOD
prints it. Methods marked not used, such as scMM, scMVP and scMSI, run
on the CPU even on a GPU host.
Some scripts use CUDA by default. For those, method_info(m)["cpu_params"]
holds the values that turn it off: scJoint use_cuda='' and scMDC
device='cpu'. Without a GPU,
mtb.run adds them
unless you set the key yourself. run(..., dry_run=True) shows the
command.
On a login node without a GPU, mtb.scan(..., assume_gpu=True) checks a
job that will run on a GPU node. The caveat of a GPU-only method then
says so, and the commands leave out the CPU flags. mtb.run_all takes assume_gpu=True
in a dry run only. In the shell, add --assume-gpu to scan or
run-all --dry-run.
RunResult¶
mtb.run returns a
RunResult. Its output is the method's primary output, for example the
embedding, ready for mtb.evaluate.
Details
A RunResult has these fields:
method: the method id.out_dir: the working directory of the run.cmd: the full command that ran, including the environment wrapper.output: the primary output, loaded. An embedding may be dims x cells.evaluatere-orients it.extra:{file: loaded output}for the method's extra outputs.stdout,stderr: the captured streams.obs_names: the cell barcodes of the output rows, in the method's cell order. It isNonewhen they could not be matched.scripts_commit: the commit of the method scripts that ran, orNonewhen they are not a git checkout.env_flavor:cpuorgpu, the build of the method's environment, elseunknown.hostname: the name of the computer that ran the method.
A non-zero exit raises RuntimeError with the end of the output. A
missing environment raises OSError before anything is written. On
Linux with no environment and no conda, the run fails with
FileNotFoundError for conda instead.
Version¶
mtb.__version__ is the installed version, and multibench --version
prints it.
Changes lists what each release adds and changes, and the old spellings with their replacements.