Skip to content

Run a method

mtb.run runs one method on your inputs and returns its output. mtb.run_all runs and scores every method that fits a dataset, as in the category tutorials.

Methods run only on Linux, each in its own environment. On macOS and Windows, mtb.run(..., dry_run=True) previews the command. Run the call itself on the Linux machine.

import multibench as mtb

mtb.data.fetch("D28")                  # 137 MB; skipped when already present
res = mtb.run(
    "SCALEX", "diagonal",
    inputs=mtb.inputs_for("D28", "diagonal", "SCALEX"),
    out_dir="runs/scalex_d28",
)
res.output.shape                       # (11014, 10): cells x latent dimensions
Details

run finds the version of the method that fits your category and files, converts the files if needed, runs the method script in out_dir and loads the output.

Relative paths in inputs and out_dir are fine.

mtb.run(..., dry_run=True) returns the command line and creates no files. multibench run ... --dry-run prints the same command, and mtb.scan(...)["command"] shows it per method.

Some commands read a file that run writes first under <out_dir>/inputs/, such as a converted AnnData or the renamed peak file of GLUE and Seurat_v3. The dry run and the caveat column of scan say so. Put multibench run in the job script for those methods.


Install the method's environment

Install the method's environment before its first run. multibench env doctor lists what is installed and what is missing.

multibench env install --methods SCALEX --packed --run
mtb.method_info("SCALEX")["env"]                # 'scmb_torch'
mtb.env.install(["SCALEX"], dry_run=False)      # the same install from Python

Three methods need more before their first run:

  • GLUE: first fetch the method scripts with multibench fetch --scripts. Then download gencode.v43.chr_patch_hapl_scaff.annotation.gtf.gz from GENCODE release 43 into <repo_path>/tools_scripts/GLUE/. multibench config prints repo_path.
  • Seurat_v5 uses your RNA and ATAC files as its paired bridge, so both must hold the same cells. Unpaired data cannot run through the package.
  • MIRA: first fetch the method scripts the same way. Then put a logger.py next to main_MIRA.py in <repo_path>/tools_scripts/MIRA/. The public scMultiBench repository does not include this file.
Details

A missing environment stops run with OSError before anything is written. On Linux the message names the install command. On macOS and Windows it says that methods run only on Linux.

A Linux host with no installed environment and no conda gives FileNotFoundError for conda instead. Run multibench env doctor to see what is missing.

--packed downloads a prebuilt environment instead of building one.

GLUE runs its human script only, so mouse data does not run through the package.

mtb.scan reports these requirements in its caveat or reason column, and a dry run prints them. mtb.method_info(m)["setup_hint"] holds the full text.


What run returns

res is a RunResult. res.output is the method's main output, here the joint embedding as an array. An embedding is a table with one row per cell. For MIRA, scMoMaT and Seurat_WNN, res.output is a cell graph instead.

res.output       # numpy array, shape (11014, 10)
res.obs_names    # barcodes of the rows of res.output
res.extra        # other declared outputs, keyed by file name ({} for SCALEX)
res.cmd          # the command line that ran
res.out_dir      # absolute path of the run folder

For a vertical folder the rows follow your cells, so adata.obsm["X_totalVI"] = res.output works when res.obs_names matches adata.obs_names.

Score the embedding with the dataset's label files, as in Evaluate a run. Give labels_for the category and the method, so the label files follow the rows of res.output.

labels = mtb.labels_for("D28", "diagonal", "SCALEX")
metrics = mtb.evaluate(res.output, labels=labels, category="diagonal",
                       metrics="all")
Details

res.method is the method id. res.stdout and res.stderr hold the method's captured output. A method that exits with an error raises RuntimeError with the end of its output.

res.obs_names is None, with a warning, when the barcodes cannot be matched to the output rows. scBridge, which reads a whole folder, is one such method. Most methods write cells x dimensions. When the first axis differs from len(res.obs_names), store res.output.T.

Without a category and a method, labels_for uses the default order: RNA before ATAC, and cty1, cty2, .... Five methods use another order. uniPort and Seurat_v5 put the ATAC cells first. StabMap in cross puts its reference batch first. SMILE and MultiVI in mosaic use their own batch order. run_all records each method's order in summary.label_order.

A method that reads only some batches gets only their label files. UINMF in cross reads batches 1 and 2 of D52, so it gets cty1 and cty2.

method_info(m)["supports"] gives each variant's output_kind, embedding or graph. Embedding metrics such as ASW do not apply to a graph. run_all scores scMoMaT through the UMAP embedding that scMoMaT also writes, and records MIRA and Seurat_WNN as RUN_OK_NO_EMBEDDING.

Each file in res.extra loads as a numpy array. A text file loads as a list of its lines.


Inputs

inputs maps each modality role to a file path or an in-memory AnnData. A role is the name of one input, such as rna or atac_gas. For a dataset folder, mtb.inputs_for builds the dict.

mtb.method_info("SCALEX")["supports"][0]["modalities"]   # ['rna', 'atac_gas']

mtb.run(
    "SCALEX", "diagonal",
    inputs={"rna": "my_rna.h5ad", "atac_gas": "my_gene_activity.h5ad"},
    out_dir="runs/scalex_mine",
)
Details

Each entry of supports is one variant. A variant is the version of a method for one category and one set of roles. Inputs whose roles match no variant raise KeyError listing the variants the method has. Rename the inputs keys to fix it.

For a method with several variants, inputs_for raises mtb.AmbiguousVariantError when the folder holds the files for none or more than one of them. Pass modalities=, such as ["rna", "adt"]. The token "atac" matches any ATAC role.

inputs_for(..., check=True) raises when a file is missing, a matrix is transposed, a label file has the wrong number of rows, or two files that must hold the same cells do not.

run converts these inputs to the .h5 files the method scripts read (see to_canonical):

  • .h5ad files
  • .csv and .tsv files (cells x features)
  • .loom files, which need pip install "multibench-sc[loom]"
  • in-memory AnnData
  • .h5 files in that layout, passed through unchanged

A MuData, in memory or as an .h5mu file, raises ValueError. Pass one modality per role, such as mdata["rna"], or write a dataset folder with mtb.io.export_dataset.

mtb.run(..., convert=False) passes your file paths to the method unchanged. It needs paths: an in-memory AnnData raises ValueError.

Your own data as a dataset folder

mtb.io.export_dataset writes an AnnData or MuData as a dataset folder, which scan, run_all and inputs_for read.

Give raw counts for RNA, ADT and peaks. The methods normalise the data themselves. The folder stores each matrix dense, so drop genes or peaks you do not need first.

With ATAC, set atac_kind= to "peak" or "gene_activity". Most methods read one of the two, and method_info(m)["atac"] names it. MultiMAP and Seurat_v3 read both. scan and run_all skip a method whose file holds the other one, also when methods= names it. mtb.run only warns, and such a method gives a wrong result.

import anndata as ad

adata = ad.read_h5ad("my_citeseq.h5ad")
# X = RNA counts, obsm["protein"] = ADT counts
mtb.io.export_dataset(adata, "data/MYCITE", rna="X", adt="obsm:protein",
                      labels="obs:celltype")
mtb.scan("MYCITE", "vertical", data_path="data", modalities=["rna", "adt"])
import mudata as md

mdata = md.read_h5mu("my_multiome.h5mu")
# labels in mdata.obs; use rna:celltype for mdata["rna"].obs
mtb.io.export_dataset(mdata, "data/MYMULTIOME", rna="rna", atac="atac",
                      atac_kind="peak", labels="obs:celltype",
                      category="vertical")
mtb.scan("MYMULTIOME", "vertical", data_path="data",
         modalities=["rna", "atac"])
# on Linux: the methods that read peaks; batch = the sample of each cell
res = mtb.run_all("MYMULTIOME", "vertical", data_path="data", out_dir="out/",
                  modalities=["rna", "atac_peak"], batch=mdata.obs["sample"])

Keep all samples in one folder and pass each cell's sample as batch=. atac_peak keeps the seven methods that read peaks. Matilda, scMDC and UnitedNet read gene activity, which goes in atac.h5 of a second folder.

import anndata as ad

rna = ad.read_h5ad("lung_rna.h5ad")
atac = ad.read_h5ad("lung_atac.h5ad")     # other cells than rna
mtb.io.export_dataset(rna, "data/LUNG", atac=atac, atac_kind="peak",
                      labels="obs:cell_type", category="diagonal")
peak_methods = mtb.find_methods("diagonal", atac="peak")
mtb.scan("LUNG", "diagonal", data_path="data", methods=peak_methods)

scan checks the four methods that read peaks. With peaks alone, only GLUE passes the file check. To add gene activity, see My ATAC is peaks only.

A mosaic or cross dataset can arrive as one file per batch. Here cite, multiome and rna_only hold one batch each.

kw = dict(labels="obs:cell_type", category="mosaic")
mtb.io.export_dataset(cite, "data/LAB", adt="obsm:protein",
                      batch_index=1, **kw)
mtb.io.export_dataset(multiome, "data/LAB", rna="rna", atac="atac",
                      atac_kind="peak", batch_index=2, **kw)
mtb.io.export_dataset(rna_only, "data/LAB", batch_index=3, **kw)
print(mtb.describe_layout("mosaic"))

batch_index=N writes the whole object as batch N. Number the batches to match a pattern the methods read. describe_layout("mosaic") lists them. For cross, write each donor file the same way with category="cross". One file with a donor column is split by batch=:

mtb.io.export_dataset(adata, "data/MYCITE3", adt="obsm:protein",
                      labels="obs:celltype", batch="obs:donor",
                      category="cross")

The cross methods read batches 1-3. describe_layout("cross") lists them.

From the shell, multibench convert takes --batch-index N.

Details: files and labels

mtb.io.to_canonical(src, "data/MYDIAG/", modality="rna") writes one modality under its file name. Keep the trailing / when the folder does not exist yet. export_dataset(..., dtype="float32") or to_canonical(..., dtype="float32") halves the size of each file. mtb.io.read_canonical reads a file back as an AnnData.

A label file is a one-column CSV: a header line x, then one label per cell, for example pd.Series(labels, name="x").to_csv(path, index=False).

A 10x Multiome AnnData read with gex_only=False holds genes and peaks in one X. Select each part with a feature filter: rna="X[feature_types=Gene Expression]" and atac="X[feature_types=Peaks]".

Running an export again raises FileExistsError for the files already in the folder. overwrite=True replaces them. A call that fails writes nothing.

mtb.scan checks the folder for each method. Its caveat column flags values that are not whole numbers. When an expected file is missing, scan and inputs_for name the file they found instead.

Details: ATAC

The ATAC file name depends on the category. Vertical reads atac.h5, which holds what method_info(m)["atac"] says. Diagonal reads atac_peak.h5 for peaks and atac_gas.h5 for gene activity. Mosaic reads atac<i>.h5, which holds peaks. The older names peak.h5, and atac.h5 for gene activity, still work.

GLUE and Seurat_v3 read peak names such as chr1:100-200. scan and run_all also skip these two methods when the peak names cannot be rewritten in that form, such as peak_0. With such names, scan cannot tell peaks from gene activity. The other methods run, with a caveat that names the file. scBridge reads the whole folder, and scan does not check its ATAC file.

To run a method that is skipped for its ATAC file, pass allow_atac_mismatch=True to scan and run_all. On the command line, the flag is --allow-atac-mismatch. run_all then runs the method, and its caveat stays in the log and in the summary.

The peak and gene-activity folders of the Multiome tab hold the same cells. To rank all their methods in one figure, give both results one dataset name:

import pandas as pd
# res_ga: the run_all result of the gene-activity folder
both = pd.concat([res.long, res_ga.long]).assign(dataset="MYMULTIOME")
fig = mtb.plot.bubble(both)

GPU or CPU?

Most methods also run without a GPU, only slower. Six need an NVIDIA GPU. On a host without one, run refuses them with OSError before starting, and scan and run_all mark them not runnable.

>>> [m for m in mtb.list_methods() if mtb.method_info(m)["requires_gpu"]]
['scBridge', 'moETM', 'SMILE', 'sciCAN', 'UnitedNet', 'iPOLNG']
>>> mtb.method_info("scJoint")["cpu_params"]
{'use_cuda': ''}
Details

scJoint and scMDC use CUDA unless told otherwise. On a host without a GPU, run adds the flags in method_info(m)["cpu_params"] and prints one line to stderr. A key you pass in params wins.

The torch environments of the tutorial methods also come as smaller CPU-only builds, which mtb.env.install picks on a host without a GPU. See Archive flavours.

A login node often has no GPU while the jobs run on GPU nodes. There, mtb.scan(..., assume_gpu=True) skips this host's GPU test, and so does mtb.run_all(..., dry_run=True, assume_gpu=True). On the command line the flag is --assume-gpu.


How the environment is entered: prefix and conda modes

run activates the installed environment folder directly, so no conda binary is needed. When no folder exists but conda is installed, it uses conda run -n <env> instead. MULTIBENCH_RUN_MODE=conda|prefix forces one mode.

Details

Environments live in <envs_dir>/<env>, where multibench env install --packed --run unpacks them. envs_dir defaults to MULTIBENCH_ENVS_DIR when set. Config lists the other defaults.

In prefix mode the command runs under bash -c after setting CONDA_PREFIX, putting <prefix>/bin first on PATH and sourcing <prefix>/etc/conda/activate.d/*.sh.

MULTIBENCH_RUN_MODE=prefix with no environment folder raises OSError naming envs_dir.


Choosing a cmd_template

Pass cmd_template to launch the command another way, for example as a Slurm job step or in a container. {env_cmd} is the command inside the method's environment. {cmd} is the bare command. With {cmd}, you manage the environment yourself.

mtb.run(..., cmd_template="srun --gres=gpu:1 {env_cmd}")
terminal
multibench run --method SCALEX --category diagonal \
    --input rna=rna.h5 --input atac_gas=atac_gas.h5 --out-dir runs/scalex \
    --runner "srun --gres=gpu:1 {env_cmd}"
mtb.run(..., cmd_template="conda run -n scmb_torch {cmd}")
mtb.run(..., cmd_template="singularity exec multibench.sif {cmd}")
# no wrapper: the active interpreter runs the script
mtb.run(..., cmd_template="{cmd}")
Details

Put {env_cmd} or {cmd} last: the rest of the template goes in front of the command. A template uses one of the two, not both.

With {env_cmd}, run checks the environment as it does without a template. With {cmd}, it skips that check.

repo_path= points at your own checkout of the scMultiBench method scripts. By default the package fetches one on the first run.


Next: turn a run's output into scIB metrics in Evaluate a run, then draw the bubble table.