Run¶
Run one method, or every method of a category, on a dataset. Methods run
only on Linux. On macOS and Windows, dry_run=True shows the command.
| Name | Summary |
|---|---|
mtb.scan |
Report what can run on a dataset, why the rest cannot, and each command. |
mtb.run |
Run one method on explicit inputs and load its output. |
mtb.run_all |
Run every runnable method on a dataset under one category and score it. |
mtb.sweep |
Run one method once per value of one hyperparameter, each in its own out_dir. |
mtb.load_batch |
Reload a saved run_all result. |
mtb.BatchResult |
The result of mtb.run_all: its summary table, long table and figure. |
mtb.scan
¶
scan(
dataset: str,
category: str | None = None,
*,
methods: list[str] | None = None,
modalities: list[str] | None = None,
data_path: Path | str | None = None,
out_dir="<out_dir>",
params: dict | None = None,
verbose: bool = True,
assume_gpu: bool = False,
allow_atac_mismatch: bool = False,
) -> DataFrame
Report what can run on a dataset, why the rest cannot, and each command.
Nothing is executed. Call it first on a new dataset;
run_all(dry_run=True) returns the same frame.
| PARAMETERS | DESCRIPTION |
|---|---|
dataset
|
Dataset folder name under
type
|
category
|
Integration category to scan;
type
|
methods
|
Method ids to include, as a list;
type
|
modalities
|
Modality tokens of one combination, e.g.
type
|
data_path
|
Data root that holds the dataset folders;
type
|
out_dir
|
Root the
type
|
params
|
type
|
verbose
|
Print one line: how many rows have their input files and their environment.
type
|
assume_gpu
|
type
|
allow_atac_mismatch
|
type
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
One row per (method, category, modalities) variant, runnable rows
first. Read |
| RAISES | DESCRIPTION |
|---|---|
FileNotFoundError
|
|
ValueError
|
Unknown |
ValueError
|
A method in |
KeyError
|
Unknown id in |
TypeError
|
|
| WARNS | DESCRIPTION |
|---|---|
UserWarning
|
|
UserWarning
|
|
Examples
>>> import multibench as mtb
>>> df = mtb.scan("D11", "vertical")
>>> df[["method", "modalities", "runnable", "reason"]]
>>> # what blocks the rest
>>> df.loc[~df.runnable, ["method", "files_reason", "env_reason"]]
>>> print(df.loc[df.files_ok, "command"].iloc[0]) # the line the run executes
>>> mtb.scan("MYCITE", "vertical", modalities=["rna", "adt"],
... data_path="/path/to/data")
Notes
Column reference. The full frame has 18 columns:
method method id
category integration category of the variant
modalities '+'-joined string ("rna+adt"); "(data_dir)" for a
directory-fed variant
env the conda env the method runs in
output_kind embedding / graph
n_tunable number of command-line hyperparameters
runtime_tier fast / medium / slow / very_slow / unknown
observed_worst_sec the slowest observed run, seconds (None = unmeasured)
caveat what the run needs besides the files, or ""
runnable both checks pass and nothing below blocks it
reason short form of the non-empty reasons, as sentences
files_ok the inputs resolve, are oriented and labelled
files_reason full file-check text, full paths
env_ok the env exists (and a GPU, when the script needs one)
env_reason full env-check text
needs_labels this variant demands a label file as an input
atac ATAC representation the method expects: 'peak' /
'gene_activity'; None when the variant takes no ATAC
command the shell line the variant would run; "" if the
inputs do not resolve
Two checks. Every row carries two independent checks, each a flag
plus a reason; runnable needs both:
files_ok/files_reason- the method's script is present, the input files resolve on disk and are oriented features x cells, every label CSV has one row per cell of the modality it labels, and adata_dirmethod (scBridge) finds the files it names. Seurat_v5'srna.h5andatac_peak.h5must hold the same cells. A diagonalatac_gas.h5must list the cells ofatac_peak.h5in its order.env_ok/env_reason- the method's conda env exists on this machine; the reason names the env and the one-method install command (multibench env install --methods X --packed --run). On macOS and Windows it says the environment runs only on Linux.env_okon a GPU-only method - when the upstream script calls CUDA unconditionally,env_okalso checks for an NVIDIA GPU (mtb.env.host_has_gpu()). Without one,env_reasongives the sentencerunwould raise:"<method> needs an NVIDIA GPU, and this computer has none. See ...".method_info(m)showsrequires_gpuand the code line ingpu_evidence.assume_gpu=True(multibench scan --assume-gpu) skips that GPU test, for a check on a login node before a GPU-node job. The row'scaveatthen saysThis check assumes the job runs on a GPU node., andcommandleaves out the CPU flags.
The file check runs whether or not any conda env is installed.
Reason columns. reason joins the non-empty reasons as sentences
and is empty only when the row is runnable. It is the short form. A
missing ATAC file leads with what the method needs and what the folder
holds (UnitedNet needs gene-activity ATAC (atac_gas.h5), and the folder
has peaks (atac_peak.h5).). Other missing files read adt.h5 is
missing.
File names are never cut. files_reason / env_reason keep the
full text with full paths; read them for a row you are debugging.
The caveat column. caveat lists what a row with files_ok
still needs, or what may go wrong without an error:
- an ATAC file that holds the other representation (a peak matrix where the method needs gene activity), or peak names the method cannot read;
- an RNA, ADT or peak file whose values are not whole numbers: the methods expect raw counts;
- diagonal: a folder whose only label file is
cty.csv; diagonal needsrna_cty.csvandatac_cty.csv; - a method that reads fewer numbered batches than the folder holds
(
UINMF reads batches 1-2 of 3. Batch 3 is not used.); - a step the user must do first, the first sentence of
method_info(m)['setup_hint'](GLUE's GENCODE annotation file); - method scripts that are not on this machine yet: the first real run
clones them with
git; on a host without network, fetch them first withmultibench fetch --scripts; - a command that reads a file
mtb.runwrites first (see the command column below).
A row given the other ATAC representation is not runnable, also when
methods= names the method, and run_all skips it. So is a GLUE or
Seurat_v3 row whose peak names are not chr:start-end.
allow_atac_mismatch=True keeps such a row runnable, with its caveat.
With MULTIBENCH_SCRIPTS_REF set to another commit than the method
scripts, or a repo_path that holds other files and no method
scripts, no row is runnable.
The command column.
commandisrun(..., dry_run=True),shlex-joined, writing under<out_dir>/<method>_<dataset>/exactly likerun_all- the literal'<out_dir>'placeholder unlessout_diris given.- Paths are absolute: a relative
out_dir, the placeholder included, is resolved against the working directory. paramsare merged in the wayrun_all(params=)merges them.- On a GPU-less host the command already carries each method's
cpu_params, the flags that turn CUDA off where a switch exists. - A row blocked only by
env_okstill shows its command. Put it in a job script once the environment is built. - Some commands read a file that
mtb.runwrites first underinputs/(Seurat_v3's renamed peak file, a converted input). Thecaveatnames the file; start such a method withmtb.runormultibench runinstead of the shell line.
The modalities column. modalities is a +-joined string here
("rna+adt"); run_all / inputs_for take a list
(["rna", "adt"]), so split on "+". The sentinel "(data_dir)"
marks a method that reads a whole folder rather than named modality
files (scBridge). For it, pass no modalities at all.
Sizing a sweep. runtime_tier / observed_worst_sec (see
method_info(m)['runtime']) let you size a sweep before launching it.
Selection and input checks.
category- a typo raisesValueErrorlisting the four valid values.methods- an unknown id raisesKeyErrorwith a did-you-mean hint; blocked rows of the selected methods are kept, with their reason. A named method with no variant undercategoryormodalitiesraisesValueError, even when other named methods have one.params- a key no variant of that method accepts raisesKeyErrornaming the accepted keys.dataset- a spelling that differs from the folder only in case ('d52'on macOS) is replaced by the on-disk spelling, with aUserWarning.
Modality tokens. modalities names one combination, in any
order. Base tokens keep a row whose modalities are exactly that
combination. Representation tokens select by what the method reads.
mtb.find_methods keeps every method that reads at least the named
modalities, so the two can list different methods.
- a list that spells a row's modalities keeps that row, with one
exception: moETM, scMM and iPOLNG read peaks through a role named
atac_gas, soatac_peakselects them andatac_gasdoes not, unlessmethods=names them; - a base token (
rna,adtor its aliasprotein,atac) matches every role of that base:atacmatchesatac,atac_gas,atac_peakand numbered roles such asatac2; - a representation token (
atac_peak/peak,atac_gas/gas/gene_activity) keeps the methods whosemethod_info(m)['atac']is that representation, likefind_methods(atac=...).atac_peakalone also keeps MultiMAP and Seurat_v3, which read both files; atac_peaktogether withatac_gaskeeps only those two;- a numbered token (
rna1) matches that role only; - an unknown token raises
ValueErrorlisting the vocabulary.
A variant that reads a whole folder (scBridge) has no modality roles.
modalities=[] selects exactly those. Other lists drop them, with a
UserWarning unless the tokens already exclude their ATAC
representation.
Choosing a category. A CITE-seq folder (rna.h5 + adt.h5 +
cty.csv) is vertical with modalities ["rna", "adt"]; RNA and
ATAC from different cells is diagonal. See mtb.list_categories
and mtb.describe_layout.
Environments and CLI. Each method runs in its own conda environment,
on Linux. List them with multibench env doctor; install one with
multibench env install --methods X --packed --run. Off Linux the
summary line says what this computer can do. multibench scan prints
a compact view by default; --columns all adds the rest, including
command.
See Also
mtb.run_all : run the runnable rows, with metrics; dry_run=True returns this frame.
mtb.describe_layout : how to lay out a dataset folder so files_ok passes.
mtb.env.doctor : the env check on its own, per env.
mtb.inputs_for : the {role: path} resolution behind files_ok.
mtb.run
¶
run(
method: str,
category: str,
*,
inputs: dict,
out_dir: str,
params: dict | None = None,
task: str | None = None,
convert: bool = True,
cmd_template: str | None = None,
repo_path: Path | None = None,
dry_run: bool = False,
)
Run one method on explicit inputs and load its output.
Use it for full control over one run; mtb.run_all runs and scores
every method of a category on a dataset.
| PARAMETERS | DESCRIPTION |
|---|---|
method
|
Method id, e.g.
type
|
category
|
Integration category of the variant:
type
|
inputs
|
type
|
out_dir
|
Folder the method writes into; created if missing.
type
|
params
|
Hyperparameter overrides merged over the variant's defaults;
type
|
task
|
Accepted for forward compatibility; currently ignored.
type
|
convert
|
Convert modality inputs to the canonical
type
|
cmd_template
|
Launcher template:
type
|
repo_path
|
Checkout holding
type
|
dry_run
|
type
|
| RETURNS | DESCRIPTION |
|---|---|
RunResult or list[str]
|
|
| RAISES | DESCRIPTION |
|---|---|
KeyError
|
Unknown method, or no variant fits |
ValueError
|
An input the run cannot convert, such as an |
ValueError
|
Input files that must hold the same cells, in one order, do not. |
OSError
|
The method needs a GPU this host lacks, or its env is not installed. |
RuntimeError
|
The method exited with a non-zero status, or its scripts could not be fetched. |
| WARNS | DESCRIPTION |
|---|---|
UserWarning
|
The input barcodes do not match the output's cells; |
UserWarning
|
An ATAC input holds the other representation or peak names the method cannot read. |
DeprecationWarning
|
UnitedNet's labels passed under the old key |
Examples
>>> import multibench as mtb
>>> inp = mtb.inputs_for("D11", "vertical", "Matilda")
>>> mtb.run("Matilda", "vertical", inputs=inp, out_dir="out/Matilda_D11",
... dry_run=True)
>>> res = mtb.run("Matilda", "vertical", inputs=inp, out_dir="out/Matilda_D11",
... params={"epochs": 20})
>>> mtb.evaluate(res.output, labels=mtb.labels_for("D11"))
>>> adata.obsm["X_Matilda"] = res.output # rows match res.obs_names
Notes
Result. RunResult carries output (the primary output, loaded),
obs_names (the input barcodes in the output's row order),
extra ({file: loaded object} for the variant's extra outputs),
cmd (the argv that ran), stdout, stderr, out_dir (an
absolute path) and method.
Dry run. dry_run=True returns the argv the real run would
execute. It uses the same variant selection, input plan, command builder
and env wrap as a real run. It creates nothing: no out_dir, no
inputs/ copies, no env check, no fetch of the method scripts.
shlex.join it for a shell line.
The preview names the files the run passes. A canonical .h5 passes
through. An AnnData or any other file becomes
<out_dir>/inputs/<role>.h5, and a peak role the method renames
becomes <out_dir>/inputs/<role>_normpeaks.h5. Input-format errors,
such as an .h5mu file, are raised by the dry run too.
The dry run prints to stderr what the real run would need first: the
method's setup_hint (method_info(m)['setup_hint']), a note when
the method scripts are not on this machine yet or not at
MULTIBENCH_SCRIPTS_REF, a note when the command reads a file under
inputs/ that the run writes first, and a note for an input path
that does not exist.
An ATAC file that mtb.scan would block gets a note too: it holds the
other representation, or peak names the method cannot read. The real
run warns and still runs.
Variant selection. Only category and the modality roles of
inputs select the variant. The modality roles are every key except the
auxiliary roles (data_dir, source_data, target_data,
out_dir) and the label roles (any key containing cty or
label). A data_dir variant declares no modalities, so category
alone selects it: pass no modality roles with it.
method_info(m)['supports'] lists every variant. A misspelt method id
raises KeyError naming the closest method id; a category and role
set with no variant raises KeyError listing the declared
(category, modalities) pairs.
Auxiliary roles (scBridge's data_dir / source_data /
target_data / source_cty / target_cty) and label files are
never converted to the canonical .h5.
Inputs. A MuData, in memory or as an .h5mu file, raises
ValueError: pass one modality per role (mdata.mod["rna"]), or
write the folder with mtb.io.export_dataset. With convert=False
every modality input must already be a file path.
Cell checks. Before anything runs, the dry run included, canonical
.h5 inputs get the cell checks of mtb.inputs_for(check=True).
Seurat_v5's rna and atac_peak must hold the same cells. A
diagonal atac_gas file must list the cells of its atac_peak
file, the one given or the one next to it, in the same order.
UnitedNet's label input is cty. The older key rna_cty still
works, with a DeprecationWarning.
GPU and CPU. cpu_params are the flags that turn CUDA off in a
script that uses it by default (method_info(m)['cpu_params']). On a
host without an NVIDIA GPU (mtb.env.host_has_gpu() is False), they
are merged into params first. A key you pass always wins.
The dry run shows these flags too; a real run prints [run] no GPU on
this host: applying <method> cpu_params {...} to stderr.
A script that calls CUDA unconditionally, with no switch
(method_info(m)['requires_gpu']), is refused on such a host with
OSError before anything is written or launched. The message is the
sentence mtb.scan reports as that row's env_reason. A dry run
still returns the argv.
Environment. With cmd_template=None the method runs in the env
mtb.method_info(method)['env'] names, entered in one of two modes,
picked per call:
prefixwhenevermtb.env.env_prefix(env)finds the env on disk: abash -cwrapper setsCONDA_PREFIX/CONDA_DEFAULT_ENV, puts<prefix>/binfirst onPATHand sources the env'sactivate.dscripts. No conda binary is needed.condaotherwise:conda run -n <env>.
MULTIBENCH_RUN_MODE=conda|prefix forces one mode. prefix with no
prefix on disk raises OSError naming envs_dir and
mtb.env.install; any other value raises ValueError.
The method process gets PYTHONNOUSERSITE=1, so user site-packages
cannot shadow the env, and MPLBACKEND=Agg unless the method sets its
own backend.
Environment variables the method itself needs are set over yours. A
value made of absolute paths, such as LD_PRELOAD, is applied only
for the paths that exist on this machine. When none does, the variable
is left as you set it, or unset, so the tool's default lookup applies.
Env check. Before any file is written, the env is looked up with
the probe mtb.scan uses. If envs are found on this machine and the
method's env is not among them, EnvironmentError (Python's alias of
OSError) is raised, naming the install command. On Linux, if the
probe finds no envs at all, the subprocess reports the failure. A
cmd_template with {cmd} takes over env control and skips the
check.
Other systems. Method environments are Linux-only. On macOS or
Windows a missing env always raises OSError, and the message starts
with that fact: run the call on a Linux machine. dry_run=True
previews the command here; the command holds this computer's paths.
Launcher templates. cmd_template wraps the command in your own
launcher. {env_cmd} is the command with the env activation above.
{cmd} is the bare command: the template must then enter an env
itself, for example "conda run -n myenv {cmd}".
A Slurm job step:
mtb.run("scMoMaT", "mosaic", inputs=mtb.inputs_for("D46", "mosaic", "scMoMaT"),
out_dir="runs/scMoMaT", cmd_template="srun --gres=gpu:1 {env_cmd}")
Request a GPU only for a method that uses one;
mtb.method_info(m)['gpu'] says which.
Paths. Relative paths in inputs and out_dir are made absolute
before the argv is built, and data_dir (like any existing directory)
gets a trailing separator, because the method runs with cwd=out_dir.
A few variants run in their script's directory instead; the output path
is passed on the command line either way.
What out_dir holds. Besides the method's own files:
- for a method fed modality files,
inputs/with their canonical.h5copies (convert=True); - a
data_dirmethod (scBridge) gets noinputs/at all.
Failures. A non-zero exit raises RuntimeError with the tail of the
method's stdout, then of its stderr (last, so a truncated message keeps
it). If the call is interrupted - Ctrl-C, or a run_all timeout - the
method's whole process tree is killed before the exception propagates.
Method scripts. The upstream scripts are never modified. A method
with a package-side driver runs the driver, which loads the unmodified
script from its own directory. The first run on a machine clones them
with git. On a host without network, fetch them first with
multibench fetch --scripts; a failed clone raises RuntimeError
naming that command.
See Also
mtb.inputs_for : builds the inputs dict from a laid-out dataset folder.
mtb.run_all : every runnable method on a dataset, scored, with failures recorded.
mtb.scan : previews the same command per method, with the file and env checks.
mtb.evaluate : scores RunResult.output.
mtb.run_all
¶
run_all(
dataset: str,
category: str,
out_dir=None,
*,
methods=None,
modalities=None,
params: dict | None = None,
data_path=None,
evaluate: bool = True,
dry_run: bool = False,
verbose: bool = True,
timeout: float | None = None,
skip_existing: bool = False,
batch=None,
assume_gpu: bool = False,
allow_atac_mismatch: bool = False,
) -> "BatchResult | pd.DataFrame"
Run every runnable method on a dataset under one category and score it.
Only rows mtb.scan marks runnable are attempted. A method's failure is
recorded, never raised, and the sweep is saved under out_dir.
| PARAMETERS | DESCRIPTION |
|---|---|
dataset
|
Dataset folder name under
type
|
category
|
Integration category:
type
|
out_dir
|
Output root, one
type
|
methods
|
Method ids to include, as a list;
type
|
modalities
|
Modality tokens of one combination, e.g.
type
|
params
|
Per-method hyperparameters,
type
|
data_path
|
Data root that holds the dataset folders;
type
|
evaluate
|
Score each embedding;
type
|
dry_run
|
type
|
verbose
|
Print
type
|
timeout
|
Per-method wall-clock cap in seconds;
type
|
skip_existing
|
Reuse an output file already in
type
|
batch
|
Batch ids, cells in the order of
type
|
assume_gpu
|
Dry run only: skip this host's GPU test, as
type
|
allow_atac_mismatch
|
type
|
| RETURNS | DESCRIPTION |
|---|---|
BatchResult or DataFrame
|
The sweep's |
| RAISES | DESCRIPTION |
|---|---|
FileNotFoundError
|
|
ValueError
|
Unknown |
ValueError
|
A |
KeyError
|
Unknown id in |
TypeError
|
A real run without |
| WARNS | DESCRIPTION |
|---|---|
UserWarning
|
|
UserWarning
|
|
UserWarning
|
A |
Examples
>>> import multibench as mtb
>>> plan = mtb.run_all("D11", "vertical", dry_run=True) # what would run?
>>> plan[["method", "modalities", "runnable", "reason"]]
>>> res = mtb.run_all("D11", "vertical", out_dir="out/", timeout=3600)
>>> res.summary # one row per method, metrics as columns
>>> res.failures # failures are recorded, not raised
Notes
Dry run. dry_run=True runs nothing and returns the
mtb.scan frame for the same selection: blocked rows are kept with
their reason, and command is rendered for out_dir (or the
literal '<out_dir>' placeholder). plan[plan.runnable] lists what
will run. len(plan) also counts blocked rows.
multibench run-all --dry-run --format csv writes the same frame.
Before the sweep. Every attempted row passed both mtb.scan checks
(input files and conda env, plus a GPU where the script needs one), so a
missing env is reported before any method starts (multibench env
doctor). Methods take minutes to hours each.
Skipped rows. A blocked row whose input files are in the folder,
or a named method with no runnable row, is logged and recorded as
SKIPPED with its reason. Other blocked rows are only counted.
Failures are recorded. In a real run a method that raises is
recorded as FAIL (with its error), one that exceeds timeout
as TIMEOUT, and the sweep moves on; a params key the variant does
not accept is a FAIL too. Check res.failures.
Timeout. Without a cap, one method that hangs stops the whole
sweep. Size it from the
runtime_tier / observed_worst_sec columns of mtb.scan (or
method_info(m)['runtime']); the slowest methods take more than 4 h.
The cap covers the run and its scoring; off the main thread it is
unavailable, with a warning.
Saved files. The result is saved automatically under out_dir
(summary.csv, failures.csv, batch_result.json, long.csv
when some method produced metrics); reload it with mtb.load_batch.
With batch=, the vector is saved as batch_<hash>.csv.
Several jobs, one folder. A later run into the same out_dir is
merged with the records already there: methods it re-ran are replaced,
the others kept. An out_dir that holds another dataset or category
raises ValueError before any method runs. Jobs running at the same
time should use one out_dir each; combine them with
multibench plot --input dir1 --input dir2.
Resuming. skip_existing=True reuses each method's existing
output. Reuse only checks that the output file exists, not
that it is complete: a method killed mid-write leaves a truncated file
that would be reused as if it had succeeded. After a hard kill, delete
that method's sub-directory before resuming.
A reused output's record copies scripts_commit, env_flavor and
hostname from the method's earlier record in out_dir. Without an
earlier record they are None, 'unknown' and ''.
Tuning. skip_existing=True together with params=... raises
ValueError on a real run: reuse is keyed on the output file, not on
params, so it would return results computed with the old parameters.
Give each setting a fresh out_dir (or leave skip_existing False),
as mtb.sweep does:
mtb.run_all("D11", "vertical", out_dir="out/lr", methods=["Multigrate"],
params={"Multigrate": {"lr": 1e-3}})
Batch vector. By default the batch metrics use the label file each
cell came from (cty1.csv -> 1 ...); batch= replaces that rule and
is recorded as batch_source='user'.
Give batch in the cell order of mtb.labels_for(dataset).
run_all puts it in each method's cell order, as it does the labels,
and keeps only the cells of the batches a method reads. A vector as long
as one method's output is used as given.
A Series or one-column DataFrame indexed by cell id is aligned to the
barcodes of the dataset's files. So is a CSV whose first column holds
them, as obs[["sample"]].to_csv(path) writes it. Ids that are not
cells of the dataset raise ValueError before any method runs.
Files without usable barcodes give a match by position, with a
UserWarning. So does a CSV whose first column holds text but none of
the barcodes. R's row numbers in that column give no warning.
Any other length marks that method RUN_OK_EVAL_FAILED (batch has N
entries, embedding has M cells); the dry run says so first. Re-score a
finished sweep with BatchResult.rescore.
ATAC files. Vertical reads atac.h5; method_info(m)["atac"]
says whether it must hold peaks or gene activity. Diagonal reads
atac_peak.h5 (peaks) and atac_gas.h5 (gene activity). Mosaic
reads atac<i>.h5 (peaks). peak.h5, and atac.h5 for gene
activity, are accepted as older names.
run_all skips a method given the other representation, also when
methods= names it. With allow_atac_mismatch=True the method runs
without an error and gives a wrong embedding.
Modality tokens. modalities follows the rule of mtb.scan.
Base tokens keep a row whose modalities are exactly that combination;
atac matches every ATAC role. Representation tokens select by what
the method reads: atac_peak / atac_gas keep the methods that
need that representation. moETM, scMM and iPOLNG read
peaks through a role named atac_gas, so atac_peak selects them
and atac_gas does not, unless methods= names them.
Errors raised.
- An unknown
category:ValueErrorlisting the four. - An unknown id in
methodsorparams:KeyErrorwith a did-you-mean hint, before anything runs. - A named method or a selection with no variant:
ValueErrorbefore anything runs, dry run included, such as "Matilda does not run on cross data." - A dry run with a
paramskey no planned variant of that method accepts:KeyErrornaming the accepted keys. - Nothing runnable:
ValueError. Its first line isNo method can run on D11 (vertical).Withmethods=, it startsNone of the requested methods (Matilda, totalVI) can run on D11 (vertical).and lists every requested variant. Withoutmethods, it gives the first 3 of N. It never lists the reasons of methods you did not ask for. On macOS or Windows, its second line says that methods run only on Linux, and it lists only the rows something else also blocks. - An
out_dirthat holds a saved result of another dataset or category:ValueError, before any method runs. skip_existing=Truewithparams, orassume_gpu=Truein a real run:ValueError; a real run checks this host's GPU.
Dataset spelling. A dataset that differs from the folder only in
case ('d52') is replaced by the on-disk spelling, with a
UserWarning, before anything is named after it.
See Also
mtb.scan : the check table this function runs from.
mtb.BatchResult : what is returned - summary, long, failures, plot, rescore.
mtb.sweep : one method over a range of one hyperparameter.
mtb.load_batch : reload a saved sweep.
mtb.run : one method, one variant, with explicit inputs.
mtb.sweep
¶
sweep(
dataset: str,
category: str,
method: str,
param: str,
values,
*,
out_dir,
modalities=None,
data_path=None,
timeout=None,
verbose: bool = True,
) -> DataFrame
Run one method once per value of one hyperparameter, each in its own out_dir.
| PARAMETERS | DESCRIPTION |
|---|---|
dataset
|
Dataset folder name under
type
|
category
|
Integration category of the variant to run.
type
|
method
|
Method id, e.g.
type
|
param
|
Hyperparameter to sweep; one of the variant's
type
|
values
|
Settings to try; each one is a separate
type
|
out_dir
|
Root folder; each setting runs under
type
|
modalities
|
Modality tokens of the variant, when the method has several in
type
|
data_path
|
Data root that holds the dataset folders;
type
|
timeout
|
Per-setting wall-clock cap in seconds, passed to
type
|
verbose
|
Print
type
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
The settings' |
| RAISES | DESCRIPTION |
|---|---|
KeyError
|
Unknown |
Examples
>>> import multibench as mtb
>>> # what can be swept
>>> mtb.params_for("Multigrate", "vertical", ["rna", "adt"])["tunable"]
>>> df = mtb.sweep("MYDATA", "vertical", "Multigrate", "lr",
... [1e-4, 1e-3, 1e-2], out_dir="out/lr")
>>> df[["lr", "status", "ARI", "NMI"]]
>>> mtb.plot.bubble(df.attrs["long"]) # one series per setting
Notes
Folder names. Each setting's folder is <param>_<value> with
. -> p and - -> m (lr=0.001 runs under
<out_dir>/lr_0p001/).
The long table. df.attrs["long"] makes each setting a separate
series ("Multigrate (lr=0.001)"), so it can go straight into
mtb.plot.bubble; .long keys rows by method, so without it every
setting would collapse onto one row. DataFrame.attrs does not
survive to_csv, so the frame is also written to
<out_dir>/sweep_long.csv (path in df.attrs["long_path"]) when any
setting produced metrics.
Failed settings. A setting that fails is not fatal: run_all
records it, so that value's row appears with status FAIL (or
TIMEOUT) and empty metrics rather than aborting the sweep. Check the
status column before reading the curve.
Untunable methods. Check mtb.params_for first: a method whose
tunable is empty hardcodes its hyperparameters upstream and cannot be
swept. sweep does not reject it up front; every setting is recorded
as FAIL.
Errors. An unknown method raises KeyError with a did-you-mean
hint; the KeyError for an unknown param lists the keys the
variant accepts. The param check needs one variant: with
modalities=None and several variants in category (Matilda under
vertical), an unknown param is recorded as FAIL for every
setting instead - pass modalities to get the KeyError. Errors of
mtb.run_all (e.g. no method can run) propagate.
See Also
mtb.params_for : the tunable hyperparameters of the variant.
mtb.run_all : what each setting runs through.
mtb.load_batch
¶
Reload a saved run_all result.
| PARAMETERS | DESCRIPTION |
|---|---|
out_dir
|
Folder holding
type
|
methods
|
Methods whose records to keep;
type
|
data_path
|
Folder that holds the dataset folder;
type
|
| RETURNS | DESCRIPTION |
|---|---|
BatchResult
|
The reloaded sweep. It remembers |
| RAISES | DESCRIPTION |
|---|---|
FileNotFoundError
|
|
ValueError
|
|
KeyError
|
A name in |
Examples
>>> import multibench as mtb
>>> res = mtb.load_batch("out/")
>>> res.summary
>>> res.plot().savefig("compare.png")
>>> mtb.load_batch(mtb.data.fetch_outputs("D11"), methods=["Matilda", "scMM"])
Notes
Files read. batch_result.json holds the per-method records.
long.csv, written when the run produced metrics, restores each
method's unrounded long table; without it BatchResult.long is rebuilt
from the records' rounded metrics.
Record order. methods= only filters: the kept records stay in the
order the tree ran them, not the order of methods.
Moved folders. A record whose out_dir does not exist is pointed
at the folder of the same name next to batch_result.json. The
dataset folder is looked up under data_path=, the recorded
data_path and then data_root, and the one found is recorded as
an absolute path. So a fetch_outputs tree or a copied run_all
folder can be re-scored.
See Also
mtb.BatchResult : the object returned.
mtb.run_all : writes the folder this function reads.
mtb.data.fetch_outputs : downloads recorded run outputs in the same layout.
mtb.BatchResult
¶
The result of mtb.run_all: its summary table, long table and figure.
mtb.run_all and mtb.load_batch build it. It keeps one record per
method, which rescore and plot read.
| PARAMETERS | DESCRIPTION |
|---|---|
records
|
One record per method run, as
type
|
dataset
|
Dataset folder name the sweep ran on.
type
|
category
|
Integration category the sweep ran under.
type
|
out_dir
|
Where the sweep wrote its outputs;
type
|
| ATTRIBUTES | DESCRIPTION |
|---|---|
records |
The raw per-method records (the same list
type
|
dataset |
Dataset folder name.
type
|
category |
Integration category.
type
|
out_dir |
The sweep's output root, or
type
|
summary |
One row per method with its status, timing and metrics (property).
type
|
long |
Long table
type
|
results |
The raw records, including every label ordering tried (property).
type
|
failures |
Methods that failed, timed out or could not be scored:
type
|
Examples
>>> import multibench as mtb
>>> res = mtb.run_all("D11", "vertical", out_dir="out/", timeout=3600)
>>> res.summary[["method", "status", "ARI", "NMI"]]
>>> res.failures # empty frame when all went well
>>> fig = res.plot(metrics=["ARI", "NMI", "ASW"])
>>> res.rescore(metrics="clustering").save("out/rescored") # nothing re-run
Notes
Size and repr. len(res) is the number of method records. The
repr counts the methods with metrics, those that ran but could not be
scored, the skipped ones and the failures.
See Also
mtb.run_all : produces one.
mtb.load_batch : reloads one from save()'s folder.
mtb.plot.bubble : the figure plot draws from long.
summary
property
¶
One row per method: status, timing, shape, label matching and every metric.
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
One row per method, sorted by |
Examples
>>> res = mtb.load_batch("out/")
>>> res.summary[["method", "status", "ARI", "label_order_confidence"]]
>>> ok = res.summary.query("status == 'CHAIN_OK'")
>>> ok.sort_values("ARI", ascending=False)
Notes
Column reference. When nothing ran the frame is empty, with the columns below except the metrics:
method method id
status outcome; see "Status values"
run_sec wall-clock seconds of the method run
output_kind embedding / graph
emb_shape [cells, dims] of the embedding; None without one
n_tunable number of command-line hyperparameters
label_order label file(s) the metrics used, in order
label_order_confidence how clearly that ordering won, 0-1
batch_source 'file_of_origin' / 'user' / None
n_batches distinct batch values used (1 = none)
ARI, NMI, ASW, ... one column per metric
label_order_note why label_order_confidence is blank
caveat scan's caveat for the method, or ""
reason why a SKIPPED method did not run, or ""
A caveat of NaN: the record was saved before this column existed;
n_batches and label_order still show which batches were read.
Status values.
CHAIN_OK- ran and scored.CHAIN_OK_GRAPH_METHOD- a graph method, scored via a secondary embedding.RUN_OK_NO_EMBEDDING- ran, but the method emits only a graph, so clustering metrics do not apply.RUN_OK_EVAL_FAILED- the method ran and produced an embedding, but scoring it failed; seeerrorinfailures.RUN_OK_NO_LABEL_MATCH- ran, but no label file matches the embedding's cell count.RUN_OK- ran withevaluate=False.TIMEOUT- exceededrun_all(timeout=...).FAIL- the method itself errored; seeerrorinfailures.SKIPPED- blocked before the run, with the reason in thereasoncolumn; infailuresonly whenmethods=named the method.
FAIL, TIMEOUT, RUN_OK_EVAL_FAILED and
RUN_OK_NO_LABEL_MATCH also appear in failures.
Graph methods. scMoMaT also writes a UMAP, which is scored:
CHAIN_OK_GRAPH_METHOD. Seurat_WNN writes only a neighbour graph:
RUN_OK_NO_EMBEDDING, and its emb_shape is None.
Batch columns. batch_source / n_batches say which batch
vector the batch metrics (ASW_batch, GC, iLISI ...) were computed
against, and how many distinct values it has (1 = none):
'file_of_origin'- each cell's label file (cty1.csv-> 1,cty2.csv-> 2 ...), the rule for multi-batch datasets.'user'- the vector passed asrun_all(batch=)/rescore(batch=).None- a single label file, so no batch structure and clustering metrics only.
Label order. label_order is which label file(s), in which
order, the metrics were computed against (e.g.
rna_cty.csv+atac_cty.csv). For diagonal data the embedding stacks
two disjoint cell sets in a method-specific order.
Label-order confidence. label_order_confidence is
(best - max(runner_up, 0)) / best over the ARI of the candidate
label orders, from 0 to 1. Near 1, one order clearly fits. Below about 0.5,
two orders scored alike; check that row's label order.
Optimistic bias. When more than one ordering is possible, the
reported metrics are those of the ordering with the highest ARI. So
they are slightly optimistic. label_order_confidence shows how far
ahead the chosen order was.
Blank confidence. The column is numeric, and a blank is NaN.
So > 0.5 is False for it and .isna() finds it. It is blank
in three cases, named by label_order_note:
"single ordering"- only one ordering was possible (normal for a paired/vertical dataset with a singlecty.csv)."winner at chance"- the winning ordering itself scored ARI < 0.05, so the ratio would compare two noise values."not scored"- the row has no metrics.
See Also
BatchResult.failures : the rows whose status means something went wrong.
BatchResult.results : the raw records with every ordering tried.
long
property
¶
The scores as a long table, for plotting.
This is what plot and mtb.plot.bubble consume.
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
Columns |
Examples
>>> res = mtb.load_batch("out/")
>>> mtb.plot.bubble(res.long, metrics=["ARI", "NMI"])
>>> res.long.pivot_table(index="method", columns="metric", values="value")
Notes
Source of the rows. Each record contributes the unrounded frame
run_all attached (or long.csv via mtb.load_batch) when
present. Otherwise the record contributes its metrics dict.
Neither that dict nor the folders that mtb.data.fetch_outputs
downloads carry scored_with.
See Also
BatchResult.plot : draws the bubble figure from this frame.
mtb.to_long : the wide -> long conversion used for the metrics dict.
results
property
¶
The raw per-method records: status, out_dir, metrics and the orderings tried.
Each record keeps the method's out_dir, which rescore reads.
| RETURNS | DESCRIPTION |
|---|---|
list[dict]
|
One dict per method. Read |
Examples
>>> res = mtb.load_batch("out/")
>>> [r["out_dir"] for r in res.results]
>>> res.results[0].get("label_order_candidates") # None with a single ordering
Notes
Record keys. A key is present only when it applies:
method, category, dataset what ran, and on what
modalities the variant's modality roles
status outcome (see BatchResult.summary)
out_dir the method's output folder
metrics {metric: value}, rounded to 4 places
params_used the hyperparameter overrides passed
run_sec, emb_shape timing and embedding shape
labels_used the label file(s) behind the metrics
label_order_candidates every label ordering tried, with its ARI
batch_source, n_batches the batch vector the batch metrics used
batch_file the file in out_dir that holds a user batch
error, traceback, note why a method failed, was skipped or was not scored
requested SKIPPED records: True when methods= named it
reused True when skip_existing reused the output
env, output_kind, n_tunable the scan row the method ran from
caveat that row's caveat, or ""
data_path, multibench_version, started_at provenance of the run
data_root data_path as an absolute path
scripts_commit, env_flavor, hostname the scripts, env build and computer
_long internal; read BatchResult.long instead
Reused outputs. A record with reused True copies
scripts_commit, env_flavor and hostname from the method's
earlier record in out_dir: the values of the run that made the
output. Without an earlier record they are None, 'unknown' and
''.
Label-order evidence. label_order_candidates holds every
ordering tried and its ARI. summary's label_order_confidence
is computed from them. It is present only when more than one ordering
was possible.
See Also
BatchResult.summary : the same records as a table.
failures
property
¶
Methods that failed, timed out or could not be scored.
run_all records failures instead of raising. Check this frame.
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
Columns |
Examples
>>> res = mtb.load_batch("out/")
>>> res.failures
>>> assert res.failures.empty, res.failures.to_string()
Notes
Statuses listed.
FAILandTIMEOUT.SKIPPED, for a method thatmethods=named and that did not run.RUN_OK_EVAL_FAILED- the embedding exists, scoring it failed.RUN_OK_NO_LABEL_MATCH- ran, but no label file has as many cells as the output, so nothing was scored. Check the folder withmtb.inputs_for(check=True)andmtb.labels_for.
Not listed. RUN_OK_NO_EMBEDDING: those methods ran correctly and
emit a graph instead of an embedding, so there is nothing for
clustering metrics to score. See summary for them.
See Also
BatchResult.summary : every method, including the ones that ran but could not be scored.
plot
¶
Bubble figure of every method that produced metrics.
Rows are methods, best first. Circle size shows the rank within a column (bigger is better). Fill compares the values in a column: the lightest is the lowest in this figure, not zero.
| PARAMETERS | DESCRIPTION |
|---|---|
**kw
|
Keyword arguments of
default
|
| RETURNS | DESCRIPTION |
|---|---|
Figure
|
Save it with |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
Nothing was scored (see |
Examples
>>> res = mtb.load_batch("out/")
>>> fig = res.plot()
>>> fig = res.plot(metrics=["ARI", "NMI", "ASW"], title="D11 vertical")
>>> fig.savefig("D11_vertical.png", dpi=200)
Notes
Keywords. metrics= sets the column order, methods= a
subset, order= the row order (unlisted methods follow best-first;
unknown names raise ValueError). There is no default title: pass
title= when one is wanted.
Reading the figure. Size and colour are relative to the methods
in this figure, so with few methods a small gap fills the whole colour
scale. Check the values in summary.
See Also
mtb.plot.bubble : the underlying function and its full keyword list.
BatchResult.long : the frame handed to it.
rescore
¶
Score the saved outputs again with new labels, batch or metrics.
No method is re-run. Each record's embedding is read back from its
out_dir.
| PARAMETERS | DESCRIPTION |
|---|---|
batch
|
Batch ids in the order of
type
|
labels
|
One cell-type label per cell; a Series or a barcode-indexed CSV is
aligned by barcode.
type
|
metrics
|
Metric family (
type
|
verbose
|
Print one line per method, and one before each label-order ranking.
type
|
| RETURNS | DESCRIPTION |
|---|---|
BatchResult
|
A new result; this one is untouched. Call |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
|
| WARNS | DESCRIPTION |
|---|---|
UserWarning
|
A Series or barcode-indexed CSV cannot be aligned and is matched by position. |
Examples
>>> res = mtb.load_batch("out/")
>>> res.rescore(metrics=["ARI", "NMI"]).summary
>>> new = res.rescore(batch="data/D11/donor.csv")
>>> new.summary[["method", "batch_source", "iLISI"]]
>>> res.rescore(labels=my_labels).save("out/rescored")
Notes
Arguments. A labels array, list or plain CSV follows the
embedding rows. Without labels=, a batch array follows the
order of mtb.labels_for(dataset). With a labels array in
embedding row order, the batch array follows the embedding rows
too. A CSV path is read like a label file. The batch is recorded as
batch_source='user'.
Aligned by barcode. A Series or one-column DataFrame with a
non-default index, or a CSV whose first column holds barcodes, is
aligned to the dataset's cells as in run_all(batch=), with the
same errors and warnings. Aligned labels go into each record's
label-order search, so label_order names the file order chosen.
Labels matched by position read (user labels).
Label order. With labels=None and several label files,
rescore ranks the file orders by ARI, which needs one Leiden sweep.
It keeps the stored order and skips the sweep when metrics= has no
ARI, NMI or iF1.
Labels without batch. Given labels and no batch, the batch
that run_all(batch=) saved is reused. Without a saved batch, every
cell is in one batch (batch_source None, n_batches 1), so
only clustering metrics are computed. Pass batch as well to get the
batch metrics.
Record status. A method that emits no embedding (graph-only) is
marked RUN_OK_NO_EMBEDDING with a note. SKIPPED, FAIL
and TIMEOUT records are kept as they are. A record whose output
file is gone becomes RUN_OK_EVAL_FAILED, with the reason in
error. So does a record whose new scoring fails, for example on a
batch of the wrong length (batch has N entries, embedding has M
cells).
Other hosts. When the dataset folder has moved, pass the folder
that now holds it as mtb.load_batch(data_path=). When the dataset
folder is not found, labels=None gives RUN_OK_NO_LABEL_MATCH
with a note. A Series or CSV is then matched by position, with a warning,
and the saved batch is not used.
Persisting. mtb.load_batch keeps returning the original result
until the new one is saved.
See Also
mtb.evaluate : the scoring function applied per record.
BatchResult.save : persist the re-scored result.
save
¶
Write this result to disk so it outlives the process.
| PARAMETERS | DESCRIPTION |
|---|---|
out_dir
|
Target folder, created if missing;
type
|
| RETURNS | DESCRIPTION |
|---|---|
Path
|
The folder written. |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
The folder holds a saved result of another dataset or category. |
Examples
>>> res = mtb.load_batch("out/")
>>> res.rescore(metrics="clustering").save("out/clustering_only")
>>> mtb.load_batch("out/clustering_only").summary
Notes
Files written.
summary.csv- thesummaryframe.long.csv- thelongframe, only when some method produced metrics.failures.csv- thefailuresframe.batch_result.json- dataset, category and the per-method records.batch_<hash>.csv- a batch vector the records were scored with.
Reload it with mtb.load_batch.
Saving into a folder that has a result. When the folder already
holds batch_result.json for the same dataset and category, the
records are merged. A method in this result replaces its earlier
record, unless it is SKIPPED and the earlier one is not; the other
earlier records are kept, and every file in the list above is
rewritten from the merged set. This result object is not changed.
A line # Merged with 1 earlier record in <folder> (StabMap).
names the kept methods.
Jobs in parallel. Jobs running at the same time should use one
out_dir each; combine them with
multibench plot --input dir1 --input dir2.
Blank confidence on disk. In summary.csv the
label_order_note column says why label_order_confidence is
empty on a row: "single ordering", "winner at chance" or "not scored".
See Also
mtb.load_batch : reads the folder back.