Skip to content

API reference

The package is imported as import multibench as mtb. The tables below list every public function by step. Each name links to its reference entry, with the signature, parameters and examples.

The five steps

A benchmark run has five steps: discover a method, prepare the data, run the method, score its output and compare the scores. Only step 3 needs Linux. The other steps also work on macOS and Windows.

import multibench as mtb

# score with the Leiden backend of the stored tables
mtb.config.DEFAULT.leiden_flavor = "leidenalg"

# Step 1: discover
mtb.find_methods(category="vertical", modalities=["rna", "adt"])

# Step 2: prepare (raw counts)
mtb.io.export_dataset(mdata, "data/MYCITE", rna="rna", adt="adt",
                      labels="rna:celltype")

# Step 3: run and score every method (Linux) ...
batch = mtb.run_all("MYCITE", "vertical", out_dir="runs/MYCITE",
                    data_path="data")

# ... or run one method
inputs = mtb.inputs_for("MYCITE", "vertical", "Matilda", data_path="data")
res = mtb.run("Matilda", "vertical", inputs=inputs, out_dir="runs/Matilda")

# Step 4: score
labels = mtb.labels_for("MYCITE", "vertical", "Matilda", data_path="data")
scores = mtb.evaluate(res.output, labels=labels)

# Step 5: compare
stored = mtb.load_results("vertical", dataset="D11")
mtb.plot.bubble(stored)
Details

Are my numbers comparable? explains the leiden_flavor line and its shell flag, --leiden-flavor leidenalg.

To place your method in a stored table, score it on the same demo dataset, such as D11 for CITE-seq. Add your own method shows the code.

mtb.to_long turns the scores into a long table, the shape mtb.load_results returns. A long table holds one score per row, with its method, dataset and metric. Concatenate the two tables and draw them with mtb.plot.bubble.

Step 1: Discover

Find the methods that fit your data, what each one needs and how to cite it. These calls work offline. A method has one variant for each category and set of modalities it supports.

entry point what it does
mtb.list_methods Method ids, optionally for one category.
mtb.find_methods Method ids that match every filter you give: category, task, modalities, ATAC representation and more.
mtb.method_info What the package knows about one method: environment, variants, runtime, GPU needs and reference.
mtb.params_for The hyperparameters a method variant accepts, with their defaults.
mtb.describe_layout How to lay out your own dataset folder for a category.
mtb.cite Citation text or BibTeX for the benchmark and the methods you ran.
mtb.list_tasks The valid task values.
mtb.list_categories The valid category values, each with a short description.

Step 2: Prepare data

Put your matrices into a dataset folder and find the files a method reads. Give raw counts: the methods normalise the data themselves. A role names one input of a method, such as rna, adt or cty.

mtb.run takes (method, category). mtb.inputs_for and mtb.labels_for take (dataset, category, method). mtb.scan takes (dataset, category). In mtb.io, mod names a MuData modality and modality names the role.

entry point what it does
mtb.io.export_dataset Write an AnnData, a MuData or loose matrices as a dataset folder.
mtb.io.to_canonical Convert one matrix to the canonical .h5 file, features x cells.
mtb.io.read_canonical Read a canonical .h5 back as an AnnData, cells x features.
mtb.io.normalize_peak_names Write a copy of a canonical .h5 with ATAC peak names spelled the way the methods expect.
mtb.inputs_for {role: path} of the files a method variant reads from a dataset folder.
mtb.labels_for {name: path} of a dataset's cell-type label files. With category and method, only the files that method reads, in the method's cell order.
mtb.data.fetch Download the named reference datasets unless they are already present.
mtb.data.fetch_outputs Download precomputed run_all outputs for a tutorial dataset.
mtb.data.fetchable The dataset ids fetch can download.

Step 3: Run

Run methods in their own environments, on Linux. mtb.scan shows what can run before you start.

entry point what it does
mtb.scan Which methods can run on a dataset, why the others cannot, and the command each would run.
mtb.run Run one method and load its output, a RunResult.
mtb.run_all Run and score every applicable method. Returns a BatchResult.
mtb.sweep Run one method once per value of one hyperparameter, each in its own out_dir.
mtb.load_batch Reload a saved run_all result.
mtb.BatchResult What run_all returns: a summary table, a long table, a figure and rescoring.
mtb.env.status Which method environments are installed on this machine.
mtb.env.plan Which environments a set of methods needs. multibench env plan also shows their sizes.
mtb.env.install Install those environments. A dry run until dry_run=False.
mtb.env.doctor Which needed environments are installed and which are missing. multibench env doctor also prints the install command.
mtb.env.recipe The build recipe of one method's environment.
mtb.env.env_prefix The folder of an installed environment, the one mtb.run enters, or None.

Step 4: Score

Compute scIB metrics for a run output.

entry point what it does
mtb.evaluate scIB metrics of a run output, scored against cell-type labels and, when given, batch labels.
mtb.to_long Reshape evaluate's result into the long results table.

Step 5: Compare

Load the stored metric tables, rank methods and draw figures.

entry point what it does
mtb.load_results The stored metric tables as one long table.
mtb.available_datasets Dataset ids with stored results.
mtb.results_coverage Which (category, dataset, method) combinations have results, and from which source.
mtb.recommend Rank methods for a category from the stored results, with the share of datasets each method was scored on.
mtb.DegenerateRerunWarning Warning class for a re-run row that most likely failed without an error.
mtb.plot.bubble Bubble table of a long results table.
mtb.plot.bar Each method's overall score across datasets, as bars.
mtb.plot.build_table The numbers behind a bubble table, without drawing it.
mtb.plot.BubbleTable What build_table returns: per-family matrices, ranks and row order.
mtb.catalog.methods The methods table as a DataFrame.
mtb.catalog.datasets The datasets table, including which datasets have stored results.
mtb.catalog.metrics The metrics table, with a description of each metric.
mtb.catalog.canonical_id The canonical method id for any known spelling.
mtb.catalog.canonical_metric The canonical metric code for any known spelling, such as "ari" -> "ARI".
mtb.catalog.known_metrics The metric codes, in family order.

Configuration and errors

The config page lists the environment variables.

entry point what it does
mtb.config.Config Paths and settings such as data_path, envs_dir and leiden_flavor. Every function reads mtb.config.DEFAULT.
mtb.AmbiguousVariantError Raised when a method has several variants and neither modalities= nor the dataset's files select one.

The same story from the shell

Each multibench subcommand wraps one Python function and takes the same argument names. multibench <command> --help lists the flags. run and run-all need Linux.

terminal
# how to lay out your data
multibench layout vertical

# write your raw counts as a dataset folder, then check it
multibench convert my.h5ad data/MYCITE --rna X --adt obsm:protein \
    --labels obs:celltype
multibench scan MYCITE --category vertical --data-path data

# the hyperparameters --param accepts
multibench params Matilda

# run one method and score it, or run and score every method
multibench run --method Matilda --category vertical --out-dir runs/Matilda \
    --input rna=data/MYCITE/rna.h5 --input adt=data/MYCITE/adt.h5 \
    --input cty=data/MYCITE/cty.csv
multibench evaluate --output runs/Matilda/embedding.h5 \
    --labels data/MYCITE/cty.csv --leiden-flavor leidenalg \
    --method Matilda --dataset MYCITE --category vertical \
    --out runs/Matilda/long.csv
multibench run-all MYCITE --category vertical --data-path data \
    --out-dir runs/MYCITE --leiden-flavor leidenalg

# your runs, then a stored table
multibench plot bubble --input runs/MYCITE --out mine.pdf
multibench plot bubble --category vertical --dataset D11 --out d11.pdf

multibench cite Matilda MOFA2
Details
  • Tables, ids and commands go to stdout. Progress, warnings and errors go to stderr, so 2>/dev/null hides them.
  • scan and run-all --dry-run clip long text to 80 characters. --columns all or --format csv|tsv|json prints every column in full, including command.
  • run --dry-run prints the command without running it. run-all --dry-run needs no --out-dir. Its commands then show <out_dir>.
  • scan --strict exits with 1 when no requested row is runnable. With --methods, it exits with 1 when any named method has no runnable row. It also exits with 1 while the method scripts are not fetched. --allow-atac-mismatch on scan and run-all counts a method as runnable when its ATAC file holds the other representation or peak names the method cannot read.
  • --assume-gpu on scan and run-all --dry-run checks a GPU-node job from a login node.
  • params METHOD lists the hyperparameters per variant. Set one with --param KEY=VALUE on run, or --param METHOD:KEY=VALUE on run-all.
  • evaluate --metrics takes a family (clustering, batch, all) or a comma-separated list of codes. Repeat --labels for several label files: they are stacked in the order given, one batch per file. --batch CSV on evaluate and run-all gives one batch id per cell.
  • plot --input together with --category adds your rows to the stored table. --dataset filters both, so your rows must be for a dataset the stored table has.
  • Pass --category to convert: the ATAC file name depends on the category.
  • The env commands are dry runs until --run.
  • config lists the paths in use and where each one comes from. info METHOD shows what one method needs. fetch downloads demo datasets, their stored outputs (--outputs) or the method scripts (--scripts).
  • The exit code is 0 on success, 1 on a runtime error, 2 on a usage error, and 3 when run-all finished but a method failed, or a method named in --methods was skipped. A line on stderr names those methods. A skipped method that --methods did not name does not set 3, and the reason column of summary.csv says why it did not run. MULTIBENCH_DEBUG=1 prints the traceback of an error.

Errors and warnings

An invalid argument raises a standard exception whose message names the valid values or the fix.

Details
  • KeyError: unknown method id, with a did-you-mean hint.
  • ValueError: unknown category, task, atac or modality token. The message lists the valid values.
  • TypeError: a single string where a list is expected, such as methods="StabMap". The message shows the list form.
  • FileNotFoundError: missing dataset folder or file. The message lists the folders that exist, or the path and the working directory.
  • FileExistsError: mtb.io.export_dataset or multibench convert would replace a file. Pass overwrite=True, or --overwrite on the command line.
  • mtb.AmbiguousVariantError: a method with several variants that neither modalities= nor the dataset's files select. It is a ValueError and also a KeyError.
  • OSError: the method's environment is not installed, the method needs an NVIDIA GPU and this computer has none, or a download of mtb.data.fetch or fetch_outputs failed. On Linux, a missing environment gets the install command. On macOS and Windows, the message says that methods run only on Linux.

The package warns with UserWarning, and with DeprecationWarning for old spellings. mtb.DegenerateRerunWarning, a UserWarning subclass, marks a re-run row that most likely failed without an error. You can filter it by class.

The metrics= vocabulary

evaluate, load_results, recommend and BatchResult.rescore take the same selector: "clustering", "batch", "all", or a list of metric codes such as ["ARI", "NMI"].

Details

The three family names select these codes:

  • "clustering": ARI, NMI, ASW, iASW, iF1, cLISI (mtb.plot.CLUSTERING_METRICS).
  • "batch": ASW_batch, GC, iLISI, kBET (mtb.plot.BATCH_METRICS).
  • "all": both families. evaluate then needs batch labels.

With the default None, evaluate computes the clustering family, plus ASW_batch, GC and iLISI when batch labels are available. load_results keeps every metric, and recommend ranks on the clustering family. evaluate computes kBET only when a list names it.

A single code still goes in a list: metrics=["ARI"]. A task name or an unknown code raises ValueError listing the valid values.

Method environments

Each method runs in a linux-64 environment, installed with mtb.env.install. Compatible methods share one environment. mtb.method_info(m)["env"] names it. On macOS and Windows, methods do not run: run and install work only as dry runs.

Details

The methods' own scripts run unmodified. The package builds the command, converts the inputs and loads the output.

mtb.run downloads the method scripts on first use. On a machine without network, run multibench fetch --scripts first. Add --ref, or set MULTIBENCH_SCRIPTS_REF, to fetch a given commit or tag. multibench config prints the commit in use, and every run records it as scripts_commit.

On macOS and Windows, mtb.env.install(..., dry_run=False) raises RuntimeError unless force=True (--force on the command line).

Run modes: prefix and conda

mtb.run enters each method's environment directly or through conda run. How the environment is entered explains both modes and MULTIBENCH_RUN_MODE=conda|prefix. A cmd_template with {cmd} replaces this, and {env_cmd} keeps it.

Archive flavours: cpu and gpu

Five torch environments have a CUDA build and a smaller CPU-only build. mtb.env.install picks the CPU build when mtb.env.host_has_gpu finds no NVIDIA GPU.

Details

The five environments are env_sciPENN, matilda, scmb_scjoint, scmb_torch and scmb_scmm2. Every other environment has a single build, the same archive for CPU and GPU hosts.

flavor="cpu" or flavor="gpu" overrides the default "auto" (--flavor on the command line). Both builds install under the environment's own name. mtb.env.status() reports the installed build as flavor: cpu, gpu, or single for an environment with one build.

GPU or CPU?

Six methods need an NVIDIA GPU. mtb.method_info(m)["requires_gpu"] says which. Without a GPU, mtb.run raises OSError before starting, and mtb.scan gives the reason. The other methods run without a GPU.

Details

The six methods are scBridge, moETM, SMILE, sciCAN, UnitedNet and iPOLNG.

method_info(m)["gpu"] says how a method uses a GPU: required, used when present, not used or unknown. multibench info METHOD prints it. Methods marked not used, such as scMM, scMVP and scMSI, run on the CPU even on a GPU host.

Some scripts use CUDA by default. For those, method_info(m)["cpu_params"] holds the values that turn it off: scJoint use_cuda='' and scMDC device='cpu'. Without a GPU, mtb.run adds them unless you set the key yourself. run(..., dry_run=True) shows the command.

On a login node without a GPU, mtb.scan(..., assume_gpu=True) checks a job that will run on a GPU node. The caveat of a GPU-only method then says so, and the commands leave out the CPU flags. mtb.run_all takes assume_gpu=True in a dry run only. In the shell, add --assume-gpu to scan or run-all --dry-run.

RunResult

mtb.run returns a RunResult. Its output is the method's primary output, for example the embedding, ready for mtb.evaluate.

Details

A RunResult has these fields:

  • method: the method id.
  • out_dir: the working directory of the run.
  • cmd: the full command that ran, including the environment wrapper.
  • output: the primary output, loaded. An embedding may be dims x cells. evaluate re-orients it.
  • extra: {file: loaded output} for the method's extra outputs.
  • stdout, stderr: the captured streams.
  • obs_names: the cell barcodes of the output rows, in the method's cell order. It is None when they could not be matched.
  • scripts_commit: the commit of the method scripts that ran, or None when they are not a git checkout.
  • env_flavor: cpu or gpu, the build of the method's environment, else unknown.
  • hostname: the name of the computer that ran the method.

A non-zero exit raises RuntimeError with the end of the output. A missing environment raises OSError before anything is written. On Linux with no environment and no conda, the run fails with FileNotFoundError for conda instead.

Version

mtb.__version__ is the installed version, and multibench --version prints it.

Changes lists what each release adds and changes, and the old spellings with their replacements.

See also