Run a method¶
mtb.run runs one
method on your inputs and returns its output. mtb.run_all runs and
scores every method that fits a dataset, as in the
category tutorials.
Methods run only on Linux, each in its own environment. On macOS and
Windows, mtb.run(..., dry_run=True) previews the command. Run the call
itself on the Linux machine.
import multibench as mtb
mtb.data.fetch("D28") # 137 MB; skipped when already present
res = mtb.run(
"SCALEX", "diagonal",
inputs=mtb.inputs_for("D28", "diagonal", "SCALEX"),
out_dir="runs/scalex_d28",
)
res.output.shape # (11014, 10): cells x latent dimensions
Details
run finds the version of the method that fits your category and
files, converts the files if needed, runs the method script in
out_dir and loads the output.
Relative paths in inputs and out_dir are fine.
mtb.run(..., dry_run=True) returns the command line and creates no
files. multibench run ... --dry-run prints the same command, and
mtb.scan(...)["command"] shows it per method.
Some commands read a file that run writes first under
<out_dir>/inputs/, such as a converted AnnData or the renamed peak
file of GLUE and Seurat_v3. The dry run and the caveat column of
scan say so. Put multibench run in the job script for those
methods.
Install the method's environment¶
Install the method's environment before its first run.
multibench env doctor lists what is installed and what is missing.
mtb.method_info("SCALEX")["env"] # 'scmb_torch'
mtb.env.install(["SCALEX"], dry_run=False) # the same install from Python
Three methods need more before their first run:
- GLUE: first fetch the method scripts with
multibench fetch --scripts. Then downloadgencode.v43.chr_patch_hapl_scaff.annotation.gtf.gzfrom GENCODE release 43 into<repo_path>/tools_scripts/GLUE/.multibench configprintsrepo_path. - Seurat_v5 uses your RNA and ATAC files as its paired bridge, so both must hold the same cells. Unpaired data cannot run through the package.
- MIRA: first fetch the method scripts the same way. Then put a
logger.pynext tomain_MIRA.pyin<repo_path>/tools_scripts/MIRA/. The public scMultiBench repository does not include this file.
Details
A missing environment stops run with OSError before anything is
written. On Linux the message names the install command. On macOS and
Windows it says that methods run only on Linux.
A Linux host with no installed environment and no conda gives
FileNotFoundError for conda instead. Run multibench env doctor
to see what is missing.
--packed downloads a prebuilt environment instead of building one.
GLUE runs its human script only, so mouse data does not run through the package.
mtb.scan reports these requirements in its caveat or reason
column, and a dry run prints them. mtb.method_info(m)["setup_hint"]
holds the full text.
What run returns¶
res is a RunResult. res.output is the method's main output, here the
joint embedding as an array. An embedding is a table with one row per cell.
For MIRA, scMoMaT and Seurat_WNN, res.output is a cell graph instead.
res.output # numpy array, shape (11014, 10)
res.obs_names # barcodes of the rows of res.output
res.extra # other declared outputs, keyed by file name ({} for SCALEX)
res.cmd # the command line that ran
res.out_dir # absolute path of the run folder
For a vertical folder the rows follow your cells, so
adata.obsm["X_totalVI"] = res.output works when res.obs_names matches
adata.obs_names.
Score the embedding with the dataset's label files, as in
Evaluate a run. Give labels_for the category and the
method, so the label files follow the rows of res.output.
labels = mtb.labels_for("D28", "diagonal", "SCALEX")
metrics = mtb.evaluate(res.output, labels=labels, category="diagonal",
metrics="all")
Details
res.method is the method id. res.stdout and res.stderr hold the
method's captured output. A method that exits with an error raises
RuntimeError with the end of its output.
res.obs_names is None, with a warning, when the barcodes cannot be
matched to the output rows. scBridge, which reads a whole folder, is
one such method. Most methods write cells x dimensions. When the
first axis differs from len(res.obs_names), store res.output.T.
Without a category and a method, labels_for uses the default order:
RNA before ATAC, and cty1, cty2, .... Five methods use another
order. uniPort and Seurat_v5 put the ATAC cells first. StabMap in
cross puts its reference batch first. SMILE and MultiVI in mosaic
use their own batch order. run_all records each method's order in
summary.label_order.
A method that reads only some batches gets only their label files.
UINMF in cross reads batches 1 and 2 of D52, so it gets cty1 and
cty2.
method_info(m)["supports"] gives each variant's output_kind,
embedding or graph. Embedding metrics such as ASW do not apply to
a graph. run_all scores scMoMaT through the UMAP embedding that
scMoMaT also writes, and records MIRA and Seurat_WNN as
RUN_OK_NO_EMBEDDING.
Each file in res.extra loads as a numpy array. A text file loads as
a list of its lines.
Inputs¶
inputs maps each modality role to a file path or an in-memory
AnnData. A role is the name of one input, such as rna or atac_gas.
For a dataset folder, mtb.inputs_for builds the dict.
mtb.method_info("SCALEX")["supports"][0]["modalities"] # ['rna', 'atac_gas']
mtb.run(
"SCALEX", "diagonal",
inputs={"rna": "my_rna.h5ad", "atac_gas": "my_gene_activity.h5ad"},
out_dir="runs/scalex_mine",
)
Details
Each entry of supports is one variant. A variant is the version of a
method for one category and one set of roles. Inputs whose roles
match no variant raise KeyError listing the variants the method
has. Rename the inputs keys to fix it.
For a method with several variants, inputs_for raises
mtb.AmbiguousVariantError when the folder holds the files for none
or more than one of them. Pass modalities=, such as
["rna", "adt"]. The token "atac" matches any ATAC role.
inputs_for(..., check=True) raises when a file is missing, a matrix
is transposed, a label file has the wrong number of rows, or two files
that must hold the same cells do not.
run converts these inputs to the .h5 files the method scripts read
(see
to_canonical):
.h5adfiles.csvand.tsvfiles (cells x features).loomfiles, which needpip install "multibench-sc[loom]"- in-memory
AnnData .h5files in that layout, passed through unchanged
A MuData, in memory or as an .h5mu file, raises ValueError. Pass
one modality per role, such as mdata["rna"], or write a dataset
folder with mtb.io.export_dataset.
mtb.run(..., convert=False) passes your file paths to the method
unchanged. It needs paths: an in-memory AnnData raises ValueError.
Your own data as a dataset folder¶
mtb.io.export_dataset writes an AnnData or MuData as a dataset folder,
which scan, run_all and inputs_for read.
Give raw counts for RNA, ADT and peaks. The methods normalise the data themselves. The folder stores each matrix dense, so drop genes or peaks you do not need first.
With ATAC, set atac_kind= to "peak" or "gene_activity". Most methods
read one of the two, and method_info(m)["atac"] names it. MultiMAP and
Seurat_v3 read both. scan and run_all skip a method whose file holds
the other one, also when methods= names it. mtb.run only warns, and
such a method gives a wrong result.
import mudata as md
mdata = md.read_h5mu("my_multiome.h5mu")
# labels in mdata.obs; use rna:celltype for mdata["rna"].obs
mtb.io.export_dataset(mdata, "data/MYMULTIOME", rna="rna", atac="atac",
atac_kind="peak", labels="obs:celltype",
category="vertical")
mtb.scan("MYMULTIOME", "vertical", data_path="data",
modalities=["rna", "atac"])
# on Linux: the methods that read peaks; batch = the sample of each cell
res = mtb.run_all("MYMULTIOME", "vertical", data_path="data", out_dir="out/",
modalities=["rna", "atac_peak"], batch=mdata.obs["sample"])
Keep all samples in one folder and pass each cell's sample as batch=.
atac_peak keeps the seven methods that read peaks. Matilda, scMDC and
UnitedNet read gene activity, which goes in atac.h5 of a second
folder.
import anndata as ad
rna = ad.read_h5ad("lung_rna.h5ad")
atac = ad.read_h5ad("lung_atac.h5ad") # other cells than rna
mtb.io.export_dataset(rna, "data/LUNG", atac=atac, atac_kind="peak",
labels="obs:cell_type", category="diagonal")
peak_methods = mtb.find_methods("diagonal", atac="peak")
mtb.scan("LUNG", "diagonal", data_path="data", methods=peak_methods)
scan checks the four methods that read peaks. With peaks alone, only
GLUE passes the file check. To add gene activity, see
My ATAC is peaks only.
A mosaic or cross dataset can arrive as one file per batch. Here
cite, multiome and rna_only hold one batch each.
kw = dict(labels="obs:cell_type", category="mosaic")
mtb.io.export_dataset(cite, "data/LAB", adt="obsm:protein",
batch_index=1, **kw)
mtb.io.export_dataset(multiome, "data/LAB", rna="rna", atac="atac",
atac_kind="peak", batch_index=2, **kw)
mtb.io.export_dataset(rna_only, "data/LAB", batch_index=3, **kw)
print(mtb.describe_layout("mosaic"))
batch_index=N writes the whole object as batch N. Number the batches
to match a pattern the methods read. describe_layout("mosaic") lists
them. For cross, write each donor file the same way with
category="cross". One file with a donor column is split by batch=:
mtb.io.export_dataset(adata, "data/MYCITE3", adt="obsm:protein",
labels="obs:celltype", batch="obs:donor",
category="cross")
The cross methods read batches 1-3. describe_layout("cross") lists
them.
From the shell, multibench convert takes --batch-index N.
Details: files and labels
mtb.io.to_canonical(src, "data/MYDIAG/", modality="rna") writes one
modality under its file name. Keep the trailing / when the folder
does not exist yet. export_dataset(..., dtype="float32") or
to_canonical(..., dtype="float32") halves the size of each file.
mtb.io.read_canonical reads a file back as an AnnData.
A label file is a one-column CSV: a header line x, then one label
per cell, for example
pd.Series(labels, name="x").to_csv(path, index=False).
A 10x Multiome AnnData read with gex_only=False holds genes and
peaks in one X. Select each part with a feature filter:
rna="X[feature_types=Gene Expression]" and
atac="X[feature_types=Peaks]".
Running an export again raises FileExistsError for the files already
in the folder. overwrite=True replaces them. A call that fails
writes nothing.
mtb.scan checks the folder for each method. Its caveat column
flags values that are not whole numbers. When an expected file is
missing, scan and inputs_for name the file they found instead.
Details: ATAC
The ATAC file name depends on the category. Vertical reads atac.h5,
which holds what method_info(m)["atac"] says. Diagonal reads
atac_peak.h5 for peaks and atac_gas.h5 for gene activity. Mosaic
reads atac<i>.h5, which holds peaks. The older names peak.h5, and
atac.h5 for gene activity, still work.
GLUE and Seurat_v3 read peak names such as chr1:100-200. scan and
run_all also skip these two methods when the peak names cannot be
rewritten in that form, such as peak_0. With such names, scan
cannot tell peaks from gene activity. The other methods run, with a
caveat that names the file. scBridge reads the whole folder, and scan
does not check its ATAC file.
To run a method that is skipped for its ATAC file, pass
allow_atac_mismatch=True to scan and run_all. On the command
line, the flag is --allow-atac-mismatch. run_all then runs the
method, and its caveat stays in the log and in the summary.
The peak and gene-activity folders of the Multiome tab hold the same cells. To rank all their methods in one figure, give both results one dataset name:
GPU or CPU?¶
Most methods also run without a GPU, only slower. Six need an NVIDIA
GPU. On a host without one, run refuses them with OSError before
starting, and scan and run_all mark them not runnable.
>>> [m for m in mtb.list_methods() if mtb.method_info(m)["requires_gpu"]]
['scBridge', 'moETM', 'SMILE', 'sciCAN', 'UnitedNet', 'iPOLNG']
>>> mtb.method_info("scJoint")["cpu_params"]
{'use_cuda': ''}
Details
scJoint and scMDC use CUDA unless told otherwise. On a host without a
GPU, run adds the flags in method_info(m)["cpu_params"] and prints
one line to stderr. A key you pass in params wins.
The torch environments of the tutorial methods also come as smaller
CPU-only builds, which mtb.env.install picks on a host without a
GPU. See Archive flavours.
A login node often has no GPU while the jobs run on GPU nodes. There,
mtb.scan(..., assume_gpu=True) skips this host's GPU test, and so
does mtb.run_all(..., dry_run=True, assume_gpu=True). On the command
line the flag is --assume-gpu.
How the environment is entered: prefix and conda modes¶
run activates the installed environment folder directly, so no conda
binary is needed. When no folder exists but conda is installed, it uses
conda run -n <env> instead. MULTIBENCH_RUN_MODE=conda|prefix forces one
mode.
Details
Environments live in <envs_dir>/<env>, where
multibench env install --packed --run unpacks them. envs_dir
defaults to MULTIBENCH_ENVS_DIR when set.
Config lists the
other defaults.
In prefix mode the command runs under bash -c after setting
CONDA_PREFIX, putting <prefix>/bin first on PATH and sourcing
<prefix>/etc/conda/activate.d/*.sh.
MULTIBENCH_RUN_MODE=prefix with no environment folder raises
OSError naming envs_dir.
Choosing a cmd_template¶
Pass cmd_template to launch the command another way, for example as a
Slurm job step or in a container. {env_cmd} is the command inside the
method's environment. {cmd} is the bare command. With {cmd}, you manage
the environment yourself.
Details
Put {env_cmd} or {cmd} last: the rest of the template goes in
front of the command. A template uses one of the two, not both.
With {env_cmd}, run checks the environment as it does without a
template. With {cmd}, it skips that check.
repo_path= points at your own checkout of the scMultiBench method
scripts. By default the package fetches one on the first run.
Next: turn a run's output into scIB metrics in Evaluate a run, then draw the bubble table.