Skip to content

Run

Run one method, or every method of a category, on a dataset. Methods run only on Linux. On macOS and Windows, dry_run=True shows the command.

Name Summary
mtb.scan Report what can run on a dataset, why the rest cannot, and each command.
mtb.run Run one method on explicit inputs and load its output.
mtb.run_all Run every runnable method on a dataset under one category and score it.
mtb.sweep Run one method once per value of one hyperparameter, each in its own out_dir.
mtb.load_batch Reload a saved run_all result.
mtb.BatchResult The result of mtb.run_all: its summary table, long table and figure.

mtb.scan

scan(
    dataset: str,
    category: str | None = None,
    *,
    methods: list[str] | None = None,
    modalities: list[str] | None = None,
    data_path: Path | str | None = None,
    out_dir="<out_dir>",
    params: dict | None = None,
    verbose: bool = True,
    assume_gpu: bool = False,
    allow_atac_mismatch: bool = False,
) -> DataFrame

Report what can run on a dataset, why the rest cannot, and each command.

Nothing is executed. Call it first on a new dataset; run_all(dry_run=True) returns the same frame.

PARAMETERS DESCRIPTION
dataset

Dataset folder name under data_path (not a path).

type str

category

Integration category to scan; None = all four.

type str | None default None

methods

Method ids to include, as a list; None = every method.

type list[str] | None default None

modalities

Modality tokens of one combination, e.g. ["rna", "adt"]; None = every combination.

type list[str] | None default None

data_path

Data root that holds the dataset folders; None = mtb.config.DEFAULT.data_path.

type Path | str | None default None

out_dir

Root the command lines write under; default the literal placeholder '<out_dir>'. Pass the real one for ready-to-run lines.

type path | str default '<out_dir>'

params

{method: {key: value}} hyperparameter overrides, rendered into command and checked against the keys each method accepts.

type dict | None default None

verbose

Print one line: how many rows have their input files and their environment.

type bool default True

assume_gpu

True = skip this host's GPU test; for a login node without a GPU that checks a GPU-node job.

type bool default False

allow_atac_mismatch

True = keep a row runnable when its ATAC file holds the other representation or peak names the method cannot read.

type bool default False

RETURNS DESCRIPTION
DataFrame

One row per (method, category, modalities) variant, runnable rows first. Read df[["method", "modalities", "runnable", "reason"]]; all 18 columns are listed in Notes.

RAISES DESCRIPTION
FileNotFoundError

<data_path>/<dataset> does not exist; the message lists the folders present.

ValueError

Unknown category or modality token.

ValueError

A method in methods has no variant in category or modalities.

KeyError

Unknown id in methods or params, or a params key no variant accepts.

TypeError

methods or modalities given as a bare string.

WARNS DESCRIPTION
UserWarning

dataset matches a folder only up to letter case.

UserWarning

modalities leaves out scBridge, which reads a folder, without excluding its ATAC form.

Examples

>>> import multibench as mtb
>>> df = mtb.scan("D11", "vertical")
>>> df[["method", "modalities", "runnable", "reason"]]
>>> # what blocks the rest
>>> df.loc[~df.runnable, ["method", "files_reason", "env_reason"]]
>>> print(df.loc[df.files_ok, "command"].iloc[0])  # the line the run executes
>>> mtb.scan("MYCITE", "vertical", modalities=["rna", "adt"],
...          data_path="/path/to/data")
Notes

Column reference. The full frame has 18 columns:

method              method id
category            integration category of the variant
modalities          '+'-joined string ("rna+adt"); "(data_dir)" for a
                    directory-fed variant
env                 the conda env the method runs in
output_kind         embedding / graph
n_tunable           number of command-line hyperparameters
runtime_tier        fast / medium / slow / very_slow / unknown
observed_worst_sec  the slowest observed run, seconds (None = unmeasured)
caveat              what the run needs besides the files, or ""
runnable            both checks pass and nothing below blocks it
reason              short form of the non-empty reasons, as sentences
files_ok            the inputs resolve, are oriented and labelled
files_reason        full file-check text, full paths
env_ok              the env exists (and a GPU, when the script needs one)
env_reason          full env-check text
needs_labels        this variant demands a label file as an input
atac                ATAC representation the method expects: 'peak' /
                    'gene_activity'; None when the variant takes no ATAC
command             the shell line the variant would run; "" if the
                    inputs do not resolve

Two checks. Every row carries two independent checks, each a flag plus a reason; runnable needs both:

  • files_ok / files_reason - the method's script is present, the input files resolve on disk and are oriented features x cells, every label CSV has one row per cell of the modality it labels, and a data_dir method (scBridge) finds the files it names. Seurat_v5's rna.h5 and atac_peak.h5 must hold the same cells. A diagonal atac_gas.h5 must list the cells of atac_peak.h5 in its order.
  • env_ok / env_reason - the method's conda env exists on this machine; the reason names the env and the one-method install command (multibench env install --methods X --packed --run). On macOS and Windows it says the environment runs only on Linux.
  • env_ok on a GPU-only method - when the upstream script calls CUDA unconditionally, env_ok also checks for an NVIDIA GPU (mtb.env.host_has_gpu()). Without one, env_reason gives the sentence run would raise: "<method> needs an NVIDIA GPU, and this computer has none. See ...". method_info(m) shows requires_gpu and the code line in gpu_evidence.
  • assume_gpu=True (multibench scan --assume-gpu) skips that GPU test, for a check on a login node before a GPU-node job. The row's caveat then says This check assumes the job runs on a GPU node., and command leaves out the CPU flags.

The file check runs whether or not any conda env is installed.

Reason columns. reason joins the non-empty reasons as sentences and is empty only when the row is runnable. It is the short form. A missing ATAC file leads with what the method needs and what the folder holds (UnitedNet needs gene-activity ATAC (atac_gas.h5), and the folder has peaks (atac_peak.h5).). Other missing files read adt.h5 is missing.

File names are never cut. files_reason / env_reason keep the full text with full paths; read them for a row you are debugging.

The caveat column. caveat lists what a row with files_ok still needs, or what may go wrong without an error:

  • an ATAC file that holds the other representation (a peak matrix where the method needs gene activity), or peak names the method cannot read;
  • an RNA, ADT or peak file whose values are not whole numbers: the methods expect raw counts;
  • diagonal: a folder whose only label file is cty.csv; diagonal needs rna_cty.csv and atac_cty.csv;
  • a method that reads fewer numbered batches than the folder holds (UINMF reads batches 1-2 of 3. Batch 3 is not used.);
  • a step the user must do first, the first sentence of method_info(m)['setup_hint'] (GLUE's GENCODE annotation file);
  • method scripts that are not on this machine yet: the first real run clones them with git; on a host without network, fetch them first with multibench fetch --scripts;
  • a command that reads a file mtb.run writes first (see the command column below).

A row given the other ATAC representation is not runnable, also when methods= names the method, and run_all skips it. So is a GLUE or Seurat_v3 row whose peak names are not chr:start-end. allow_atac_mismatch=True keeps such a row runnable, with its caveat. With MULTIBENCH_SCRIPTS_REF set to another commit than the method scripts, or a repo_path that holds other files and no method scripts, no row is runnable.

The command column.

  • command is run(..., dry_run=True), shlex-joined, writing under <out_dir>/<method>_<dataset>/ exactly like run_all - the literal '<out_dir>' placeholder unless out_dir is given.
  • Paths are absolute: a relative out_dir, the placeholder included, is resolved against the working directory.
  • params are merged in the way run_all(params=) merges them.
  • On a GPU-less host the command already carries each method's cpu_params, the flags that turn CUDA off where a switch exists.
  • A row blocked only by env_ok still shows its command. Put it in a job script once the environment is built.
  • Some commands read a file that mtb.run writes first under inputs/ (Seurat_v3's renamed peak file, a converted input). The caveat names the file; start such a method with mtb.run or multibench run instead of the shell line.

The modalities column. modalities is a +-joined string here ("rna+adt"); run_all / inputs_for take a list (["rna", "adt"]), so split on "+". The sentinel "(data_dir)" marks a method that reads a whole folder rather than named modality files (scBridge). For it, pass no modalities at all.

Sizing a sweep. runtime_tier / observed_worst_sec (see method_info(m)['runtime']) let you size a sweep before launching it.

Selection and input checks.

  • category - a typo raises ValueError listing the four valid values.
  • methods - an unknown id raises KeyError with a did-you-mean hint; blocked rows of the selected methods are kept, with their reason. A named method with no variant under category or modalities raises ValueError, even when other named methods have one.
  • params - a key no variant of that method accepts raises KeyError naming the accepted keys.
  • dataset - a spelling that differs from the folder only in case ('d52' on macOS) is replaced by the on-disk spelling, with a UserWarning.

Modality tokens. modalities names one combination, in any order. Base tokens keep a row whose modalities are exactly that combination. Representation tokens select by what the method reads. mtb.find_methods keeps every method that reads at least the named modalities, so the two can list different methods.

  • a list that spells a row's modalities keeps that row, with one exception: moETM, scMM and iPOLNG read peaks through a role named atac_gas, so atac_peak selects them and atac_gas does not, unless methods= names them;
  • a base token (rna, adt or its alias protein, atac) matches every role of that base: atac matches atac, atac_gas, atac_peak and numbered roles such as atac2;
  • a representation token (atac_peak / peak, atac_gas / gas / gene_activity) keeps the methods whose method_info(m)['atac'] is that representation, like find_methods(atac=...). atac_peak alone also keeps MultiMAP and Seurat_v3, which read both files;
  • atac_peak together with atac_gas keeps only those two;
  • a numbered token (rna1) matches that role only;
  • an unknown token raises ValueError listing the vocabulary.

A variant that reads a whole folder (scBridge) has no modality roles. modalities=[] selects exactly those. Other lists drop them, with a UserWarning unless the tokens already exclude their ATAC representation.

Choosing a category. A CITE-seq folder (rna.h5 + adt.h5 + cty.csv) is vertical with modalities ["rna", "adt"]; RNA and ATAC from different cells is diagonal. See mtb.list_categories and mtb.describe_layout.

Environments and CLI. Each method runs in its own conda environment, on Linux. List them with multibench env doctor; install one with multibench env install --methods X --packed --run. Off Linux the summary line says what this computer can do. multibench scan prints a compact view by default; --columns all adds the rest, including command.

See Also

mtb.run_all : run the runnable rows, with metrics; dry_run=True returns this frame.

mtb.describe_layout : how to lay out a dataset folder so files_ok passes.

mtb.env.doctor : the env check on its own, per env.

mtb.inputs_for : the {role: path} resolution behind files_ok.

mtb.run

run(
    method: str,
    category: str,
    *,
    inputs: dict,
    out_dir: str,
    params: dict | None = None,
    task: str | None = None,
    convert: bool = True,
    cmd_template: str | None = None,
    repo_path: Path | None = None,
    dry_run: bool = False,
)

Run one method on explicit inputs and load its output.

Use it for full control over one run; mtb.run_all runs and scores every method of a category on a dataset.

PARAMETERS DESCRIPTION
method

Method id, e.g. "Matilda"; see mtb.list_methods().

type str

category

Integration category of the variant: vertical, diagonal, mosaic or cross.

type str

inputs

{role: path or AnnData}, usually from mtb.inputs_for; its modality roles select the variant.

type dict

out_dir

Folder the method writes into; created if missing.

type str

params

Hyperparameter overrides merged over the variant's defaults; None = the defaults.

type dict | None default None

task

Accepted for forward compatibility; currently ignored.

type str | None default None

convert

Convert modality inputs to the canonical .h5 layout before the run.

type bool default True

cmd_template

Launcher template: {cmd} = the bare command, {env_cmd} = the command inside the method env; None = enter the method env.

type str | None default None

repo_path

Checkout holding tools_scripts/; None = mtb.config.DEFAULT.repo_path if it has one, else the package root if it has one, else a clone into the former.

type Path | None default None

dry_run

True = return the command without running it or writing anything.

type bool default False

RETURNS DESCRIPTION
RunResult or list[str]

RunResult - read output (the primary output, loaded), obs_names (its cell barcodes) and stderr. With dry_run=True, the argv list.

RAISES DESCRIPTION
KeyError

Unknown method, or no variant fits category and the input roles.

ValueError

An input the run cannot convert, such as an .h5mu file or a MuData.

ValueError

Input files that must hold the same cells, in one order, do not.

OSError

The method needs a GPU this host lacks, or its env is not installed.

RuntimeError

The method exited with a non-zero status, or its scripts could not be fetched.

WARNS DESCRIPTION
UserWarning

The input barcodes do not match the output's cells; obs_names is None.

UserWarning

An ATAC input holds the other representation or peak names the method cannot read.

DeprecationWarning

UnitedNet's labels passed under the old key rna_cty; the key is cty.

Examples

>>> import multibench as mtb
>>> inp = mtb.inputs_for("D11", "vertical", "Matilda")
>>> mtb.run("Matilda", "vertical", inputs=inp, out_dir="out/Matilda_D11",
...         dry_run=True)
>>> res = mtb.run("Matilda", "vertical", inputs=inp, out_dir="out/Matilda_D11",
...               params={"epochs": 20})
>>> mtb.evaluate(res.output, labels=mtb.labels_for("D11"))
>>> adata.obsm["X_Matilda"] = res.output      # rows match res.obs_names
Notes

Result. RunResult carries output (the primary output, loaded), obs_names (the input barcodes in the output's row order), extra ({file: loaded object} for the variant's extra outputs), cmd (the argv that ran), stdout, stderr, out_dir (an absolute path) and method.

Dry run. dry_run=True returns the argv the real run would execute. It uses the same variant selection, input plan, command builder and env wrap as a real run. It creates nothing: no out_dir, no inputs/ copies, no env check, no fetch of the method scripts. shlex.join it for a shell line.

The preview names the files the run passes. A canonical .h5 passes through. An AnnData or any other file becomes <out_dir>/inputs/<role>.h5, and a peak role the method renames becomes <out_dir>/inputs/<role>_normpeaks.h5. Input-format errors, such as an .h5mu file, are raised by the dry run too.

The dry run prints to stderr what the real run would need first: the method's setup_hint (method_info(m)['setup_hint']), a note when the method scripts are not on this machine yet or not at MULTIBENCH_SCRIPTS_REF, a note when the command reads a file under inputs/ that the run writes first, and a note for an input path that does not exist.

An ATAC file that mtb.scan would block gets a note too: it holds the other representation, or peak names the method cannot read. The real run warns and still runs.

Variant selection. Only category and the modality roles of inputs select the variant. The modality roles are every key except the auxiliary roles (data_dir, source_data, target_data, out_dir) and the label roles (any key containing cty or label). A data_dir variant declares no modalities, so category alone selects it: pass no modality roles with it.

method_info(m)['supports'] lists every variant. A misspelt method id raises KeyError naming the closest method id; a category and role set with no variant raises KeyError listing the declared (category, modalities) pairs.

Auxiliary roles (scBridge's data_dir / source_data / target_data / source_cty / target_cty) and label files are never converted to the canonical .h5.

Inputs. A MuData, in memory or as an .h5mu file, raises ValueError: pass one modality per role (mdata.mod["rna"]), or write the folder with mtb.io.export_dataset. With convert=False every modality input must already be a file path.

Cell checks. Before anything runs, the dry run included, canonical .h5 inputs get the cell checks of mtb.inputs_for(check=True). Seurat_v5's rna and atac_peak must hold the same cells. A diagonal atac_gas file must list the cells of its atac_peak file, the one given or the one next to it, in the same order.

UnitedNet's label input is cty. The older key rna_cty still works, with a DeprecationWarning.

GPU and CPU. cpu_params are the flags that turn CUDA off in a script that uses it by default (method_info(m)['cpu_params']). On a host without an NVIDIA GPU (mtb.env.host_has_gpu() is False), they are merged into params first. A key you pass always wins.

The dry run shows these flags too; a real run prints [run] no GPU on this host: applying <method> cpu_params {...} to stderr.

A script that calls CUDA unconditionally, with no switch (method_info(m)['requires_gpu']), is refused on such a host with OSError before anything is written or launched. The message is the sentence mtb.scan reports as that row's env_reason. A dry run still returns the argv.

Environment. With cmd_template=None the method runs in the env mtb.method_info(method)['env'] names, entered in one of two modes, picked per call:

  • prefix whenever mtb.env.env_prefix(env) finds the env on disk: a bash -c wrapper sets CONDA_PREFIX / CONDA_DEFAULT_ENV, puts <prefix>/bin first on PATH and sources the env's activate.d scripts. No conda binary is needed.
  • conda otherwise: conda run -n <env>.

MULTIBENCH_RUN_MODE=conda|prefix forces one mode. prefix with no prefix on disk raises OSError naming envs_dir and mtb.env.install; any other value raises ValueError.

The method process gets PYTHONNOUSERSITE=1, so user site-packages cannot shadow the env, and MPLBACKEND=Agg unless the method sets its own backend.

Environment variables the method itself needs are set over yours. A value made of absolute paths, such as LD_PRELOAD, is applied only for the paths that exist on this machine. When none does, the variable is left as you set it, or unset, so the tool's default lookup applies.

Env check. Before any file is written, the env is looked up with the probe mtb.scan uses. If envs are found on this machine and the method's env is not among them, EnvironmentError (Python's alias of OSError) is raised, naming the install command. On Linux, if the probe finds no envs at all, the subprocess reports the failure. A cmd_template with {cmd} takes over env control and skips the check.

Other systems. Method environments are Linux-only. On macOS or Windows a missing env always raises OSError, and the message starts with that fact: run the call on a Linux machine. dry_run=True previews the command here; the command holds this computer's paths.

Launcher templates. cmd_template wraps the command in your own launcher. {env_cmd} is the command with the env activation above. {cmd} is the bare command: the template must then enter an env itself, for example "conda run -n myenv {cmd}".

A Slurm job step:

mtb.run("scMoMaT", "mosaic", inputs=mtb.inputs_for("D46", "mosaic", "scMoMaT"),
        out_dir="runs/scMoMaT", cmd_template="srun --gres=gpu:1 {env_cmd}")

Request a GPU only for a method that uses one; mtb.method_info(m)['gpu'] says which.

Paths. Relative paths in inputs and out_dir are made absolute before the argv is built, and data_dir (like any existing directory) gets a trailing separator, because the method runs with cwd=out_dir. A few variants run in their script's directory instead; the output path is passed on the command line either way.

What out_dir holds. Besides the method's own files:

  • for a method fed modality files, inputs/ with their canonical .h5 copies (convert=True);
  • a data_dir method (scBridge) gets no inputs/ at all.

Failures. A non-zero exit raises RuntimeError with the tail of the method's stdout, then of its stderr (last, so a truncated message keeps it). If the call is interrupted - Ctrl-C, or a run_all timeout - the method's whole process tree is killed before the exception propagates.

Method scripts. The upstream scripts are never modified. A method with a package-side driver runs the driver, which loads the unmodified script from its own directory. The first run on a machine clones them with git. On a host without network, fetch them first with multibench fetch --scripts; a failed clone raises RuntimeError naming that command.

See Also

mtb.inputs_for : builds the inputs dict from a laid-out dataset folder.

mtb.run_all : every runnable method on a dataset, scored, with failures recorded.

mtb.scan : previews the same command per method, with the file and env checks.

mtb.evaluate : scores RunResult.output.

mtb.run_all

run_all(
    dataset: str,
    category: str,
    out_dir=None,
    *,
    methods=None,
    modalities=None,
    params: dict | None = None,
    data_path=None,
    evaluate: bool = True,
    dry_run: bool = False,
    verbose: bool = True,
    timeout: float | None = None,
    skip_existing: bool = False,
    batch=None,
    assume_gpu: bool = False,
    allow_atac_mismatch: bool = False,
) -> "BatchResult | pd.DataFrame"

Run every runnable method on a dataset under one category and score it.

Only rows mtb.scan marks runnable are attempted. A method's failure is recorded, never raised, and the sweep is saved under out_dir.

PARAMETERS DESCRIPTION
dataset

Dataset folder name under data_path, e.g. "MYCITE" (not a path).

type str

category

Integration category: vertical, diagonal, mosaic or cross.

type str

out_dir

Output root, one <out_dir>/<method>_<dataset>/ per method; required unless dry_run=True.

type path | None default None

methods

Method ids to include, as a list; None = every runnable method.

type list[str] | None default None

modalities

Modality tokens of one combination, e.g. ["rna", "adt"]; None = every combination.

type list[str] | None default None

params

Per-method hyperparameters, {"Cobolt": {"lr": 1e-3}}; see mtb.params_for for the accepted keys.

type dict | None default None

data_path

Data root that holds the dataset folders; None = mtb.config.DEFAULT.data_path.

type path | None default None

evaluate

Score each embedding; False only runs (status RUN_OK).

type bool default True

dry_run

True = return the mtb.scan frame for this selection and run nothing.

type bool default False

verbose

Print [run_all] ... progress lines.

type bool default True

timeout

Per-method wall-clock cap in seconds; None = no cap.

type float | None default None

skip_existing

Reuse an output file already in out_dir instead of re-running the method, to resume an interrupted sweep.

type bool default False

batch

Batch ids, cells in the order of mtb.labels_for(dataset); a Series or a barcode-indexed CSV is aligned by barcode. None = each cell's label file.

type array - like | Series | path | None default None

assume_gpu

Dry run only: skip this host's GPU test, as mtb.scan(assume_gpu=True) does.

type bool default False

allow_atac_mismatch

True = also run a method whose ATAC file holds the other representation or unreadable peak names.

type bool default False

RETURNS DESCRIPTION
BatchResult or DataFrame

The sweep's BatchResult; read summary and failures first. With dry_run=True, the mtb.scan frame.

RAISES DESCRIPTION
FileNotFoundError

<data_path>/<dataset> does not exist; the message lists the folders present.

ValueError

Unknown category, no matching variant, nothing runnable, a mismatched out_dir, or conflicting arguments (Notes).

ValueError

A batch Series or CSV holds ids that are not cells of the dataset.

KeyError

Unknown id in methods or params; on a dry run, a rejected params key.

TypeError

A real run without out_dir; methods or modalities given as a bare string.

WARNS DESCRIPTION
UserWarning

dataset matches a folder only up to letter case.

UserWarning

modalities leaves out scBridge, which reads a folder, without excluding its ATAC form.

UserWarning

A batch Series or barcode-indexed CSV cannot be aligned and is matched by position.

Examples

>>> import multibench as mtb
>>> plan = mtb.run_all("D11", "vertical", dry_run=True)  # what would run?
>>> plan[["method", "modalities", "runnable", "reason"]]
>>> res = mtb.run_all("D11", "vertical", out_dir="out/", timeout=3600)
>>> res.summary        # one row per method, metrics as columns
>>> res.failures       # failures are recorded, not raised
Notes

Dry run. dry_run=True runs nothing and returns the mtb.scan frame for the same selection: blocked rows are kept with their reason, and command is rendered for out_dir (or the literal '<out_dir>' placeholder). plan[plan.runnable] lists what will run. len(plan) also counts blocked rows. multibench run-all --dry-run --format csv writes the same frame.

Before the sweep. Every attempted row passed both mtb.scan checks (input files and conda env, plus a GPU where the script needs one), so a missing env is reported before any method starts (multibench env doctor). Methods take minutes to hours each.

Skipped rows. A blocked row whose input files are in the folder, or a named method with no runnable row, is logged and recorded as SKIPPED with its reason. Other blocked rows are only counted.

Failures are recorded. In a real run a method that raises is recorded as FAIL (with its error), one that exceeds timeout as TIMEOUT, and the sweep moves on; a params key the variant does not accept is a FAIL too. Check res.failures.

Timeout. Without a cap, one method that hangs stops the whole sweep. Size it from the runtime_tier / observed_worst_sec columns of mtb.scan (or method_info(m)['runtime']); the slowest methods take more than 4 h. The cap covers the run and its scoring; off the main thread it is unavailable, with a warning.

Saved files. The result is saved automatically under out_dir (summary.csv, failures.csv, batch_result.json, long.csv when some method produced metrics); reload it with mtb.load_batch. With batch=, the vector is saved as batch_<hash>.csv.

Several jobs, one folder. A later run into the same out_dir is merged with the records already there: methods it re-ran are replaced, the others kept. An out_dir that holds another dataset or category raises ValueError before any method runs. Jobs running at the same time should use one out_dir each; combine them with multibench plot --input dir1 --input dir2.

Resuming. skip_existing=True reuses each method's existing output. Reuse only checks that the output file exists, not that it is complete: a method killed mid-write leaves a truncated file that would be reused as if it had succeeded. After a hard kill, delete that method's sub-directory before resuming.

A reused output's record copies scripts_commit, env_flavor and hostname from the method's earlier record in out_dir. Without an earlier record they are None, 'unknown' and ''.

Tuning. skip_existing=True together with params=... raises ValueError on a real run: reuse is keyed on the output file, not on params, so it would return results computed with the old parameters. Give each setting a fresh out_dir (or leave skip_existing False), as mtb.sweep does:

mtb.run_all("D11", "vertical", out_dir="out/lr", methods=["Multigrate"],
            params={"Multigrate": {"lr": 1e-3}})

Batch vector. By default the batch metrics use the label file each cell came from (cty1.csv -> 1 ...); batch= replaces that rule and is recorded as batch_source='user'.

Give batch in the cell order of mtb.labels_for(dataset). run_all puts it in each method's cell order, as it does the labels, and keeps only the cells of the batches a method reads. A vector as long as one method's output is used as given.

A Series or one-column DataFrame indexed by cell id is aligned to the barcodes of the dataset's files. So is a CSV whose first column holds them, as obs[["sample"]].to_csv(path) writes it. Ids that are not cells of the dataset raise ValueError before any method runs.

Files without usable barcodes give a match by position, with a UserWarning. So does a CSV whose first column holds text but none of the barcodes. R's row numbers in that column give no warning.

Any other length marks that method RUN_OK_EVAL_FAILED (batch has N entries, embedding has M cells); the dry run says so first. Re-score a finished sweep with BatchResult.rescore.

ATAC files. Vertical reads atac.h5; method_info(m)["atac"] says whether it must hold peaks or gene activity. Diagonal reads atac_peak.h5 (peaks) and atac_gas.h5 (gene activity). Mosaic reads atac<i>.h5 (peaks). peak.h5, and atac.h5 for gene activity, are accepted as older names.

run_all skips a method given the other representation, also when methods= names it. With allow_atac_mismatch=True the method runs without an error and gives a wrong embedding.

Modality tokens. modalities follows the rule of mtb.scan. Base tokens keep a row whose modalities are exactly that combination; atac matches every ATAC role. Representation tokens select by what the method reads: atac_peak / atac_gas keep the methods that need that representation. moETM, scMM and iPOLNG read peaks through a role named atac_gas, so atac_peak selects them and atac_gas does not, unless methods= names them.

Errors raised.

  • An unknown category: ValueError listing the four.
  • An unknown id in methods or params: KeyError with a did-you-mean hint, before anything runs.
  • A named method or a selection with no variant: ValueError before anything runs, dry run included, such as "Matilda does not run on cross data."
  • A dry run with a params key no planned variant of that method accepts: KeyError naming the accepted keys.
  • Nothing runnable: ValueError. Its first line is No method can run on D11 (vertical). With methods=, it starts None of the requested methods (Matilda, totalVI) can run on D11 (vertical). and lists every requested variant. Without methods, it gives the first 3 of N. It never lists the reasons of methods you did not ask for. On macOS or Windows, its second line says that methods run only on Linux, and it lists only the rows something else also blocks.
  • An out_dir that holds a saved result of another dataset or category: ValueError, before any method runs.
  • skip_existing=True with params, or assume_gpu=True in a real run: ValueError; a real run checks this host's GPU.

Dataset spelling. A dataset that differs from the folder only in case ('d52') is replaced by the on-disk spelling, with a UserWarning, before anything is named after it.

See Also

mtb.scan : the check table this function runs from.

mtb.BatchResult : what is returned - summary, long, failures, plot, rescore.

mtb.sweep : one method over a range of one hyperparameter.

mtb.load_batch : reload a saved sweep.

mtb.run : one method, one variant, with explicit inputs.

mtb.sweep

sweep(
    dataset: str,
    category: str,
    method: str,
    param: str,
    values,
    *,
    out_dir,
    modalities=None,
    data_path=None,
    timeout=None,
    verbose: bool = True,
) -> DataFrame

Run one method once per value of one hyperparameter, each in its own out_dir.

PARAMETERS DESCRIPTION
dataset

Dataset folder name under data_path, as for mtb.run_all.

type str

category

Integration category of the variant to run.

type str

method

Method id, e.g. "Matilda"; see mtb.list_methods().

type str

param

Hyperparameter to sweep; one of the variant's tunable keys (mtb.params_for).

type str

values

Settings to try; each one is a separate mtb.run_all.

type iterable

out_dir

Root folder; each setting runs under <out_dir>/<param>_<value>/.

type path

modalities

Modality tokens of the variant, when the method has several in category; None = every variant.

type list[str] | None default None

data_path

Data root that holds the dataset folders; None = mtb.config.DEFAULT.data_path.

type path | None default None

timeout

Per-setting wall-clock cap in seconds, passed to run_all; None = no cap.

type float | None default None

verbose

Print run_all's progress lines.

type bool default True

RETURNS DESCRIPTION
DataFrame

The settings' BatchResult.summary rows stacked, the swept value first (column named param). A long table for plotting is in df.attrs["long"].

RAISES DESCRIPTION
KeyError

Unknown method, or param not among the tunable keys of a single variant.

Examples

>>> import multibench as mtb
>>> # what can be swept
>>> mtb.params_for("Multigrate", "vertical", ["rna", "adt"])["tunable"]
>>> df = mtb.sweep("MYDATA", "vertical", "Multigrate", "lr",
...                [1e-4, 1e-3, 1e-2], out_dir="out/lr")
>>> df[["lr", "status", "ARI", "NMI"]]
>>> mtb.plot.bubble(df.attrs["long"])      # one series per setting
Notes

Folder names. Each setting's folder is <param>_<value> with . -> p and - -> m (lr=0.001 runs under <out_dir>/lr_0p001/).

The long table. df.attrs["long"] makes each setting a separate series ("Multigrate (lr=0.001)"), so it can go straight into mtb.plot.bubble; .long keys rows by method, so without it every setting would collapse onto one row. DataFrame.attrs does not survive to_csv, so the frame is also written to <out_dir>/sweep_long.csv (path in df.attrs["long_path"]) when any setting produced metrics.

Failed settings. A setting that fails is not fatal: run_all records it, so that value's row appears with status FAIL (or TIMEOUT) and empty metrics rather than aborting the sweep. Check the status column before reading the curve.

Untunable methods. Check mtb.params_for first: a method whose tunable is empty hardcodes its hyperparameters upstream and cannot be swept. sweep does not reject it up front; every setting is recorded as FAIL.

Errors. An unknown method raises KeyError with a did-you-mean hint; the KeyError for an unknown param lists the keys the variant accepts. The param check needs one variant: with modalities=None and several variants in category (Matilda under vertical), an unknown param is recorded as FAIL for every setting instead - pass modalities to get the KeyError. Errors of mtb.run_all (e.g. no method can run) propagate.

See Also

mtb.params_for : the tunable hyperparameters of the variant.

mtb.run_all : what each setting runs through.

mtb.load_batch

load_batch(out_dir, *, methods=None, data_path=None) -> 'BatchResult'

Reload a saved run_all result.

PARAMETERS DESCRIPTION
out_dir

Folder holding batch_result.json: a run_all out_dir or a mtb.data.fetch_outputs tree.

type path - like

methods

Methods whose records to keep; None = every record.

type list[str] | None default None

data_path

Folder that holds the dataset folder; None = the path each record saved.

type path - like | None default None

RETURNS DESCRIPTION
BatchResult

The reloaded sweep. It remembers out_dir, so save() with no argument writes back to the same folder.

RAISES DESCRIPTION
FileNotFoundError

out_dir holds no batch_result.json.

ValueError

data_path does not hold the dataset folder.

KeyError

A name in methods has no record; the message lists the methods that do.

Examples

>>> import multibench as mtb
>>> res = mtb.load_batch("out/")
>>> res.summary
>>> res.plot().savefig("compare.png")
>>> mtb.load_batch(mtb.data.fetch_outputs("D11"), methods=["Matilda", "scMM"])
Notes

Files read. batch_result.json holds the per-method records. long.csv, written when the run produced metrics, restores each method's unrounded long table; without it BatchResult.long is rebuilt from the records' rounded metrics.

Record order. methods= only filters: the kept records stay in the order the tree ran them, not the order of methods.

Moved folders. A record whose out_dir does not exist is pointed at the folder of the same name next to batch_result.json. The dataset folder is looked up under data_path=, the recorded data_path and then data_root, and the one found is recorded as an absolute path. So a fetch_outputs tree or a copied run_all folder can be re-scored.

See Also

mtb.BatchResult : the object returned.

mtb.run_all : writes the folder this function reads.

mtb.data.fetch_outputs : downloads recorded run outputs in the same layout.

mtb.BatchResult

BatchResult(records, dataset, category, out_dir=None)

The result of mtb.run_all: its summary table, long table and figure.

mtb.run_all and mtb.load_batch build it. It keeps one record per method, which rescore and plot read.

PARAMETERS DESCRIPTION
records

One record per method run, as run_all builds them.

type list[dict]

dataset

Dataset folder name the sweep ran on.

type str

category

Integration category the sweep ran under.

type str

out_dir

Where the sweep wrote its outputs; save() defaults to it.

type path | None default None

ATTRIBUTES DESCRIPTION
records

The raw per-method records (the same list results returns).

type list[dict]

dataset

Dataset folder name.

type str

category

Integration category.

type str

out_dir

The sweep's output root, or None for an in-memory result.

type path | None

summary

One row per method with its status, timing and metrics (property).

type DataFrame

long

Long table metric, value, method, dataset, category, ... for plotting (property).

type DataFrame

results

The raw records, including every label ordering tried (property).

type list[dict]

failures

Methods that failed, timed out or could not be scored: method, status, error (property).

type DataFrame

Examples

>>> import multibench as mtb
>>> res = mtb.run_all("D11", "vertical", out_dir="out/", timeout=3600)
>>> res.summary[["method", "status", "ARI", "NMI"]]
>>> res.failures                              # empty frame when all went well
>>> fig = res.plot(metrics=["ARI", "NMI", "ASW"])
>>> res.rescore(metrics="clustering").save("out/rescored")   # nothing re-run
Notes

Size and repr. len(res) is the number of method records. The repr counts the methods with metrics, those that ran but could not be scored, the skipped ones and the failures.

See Also

mtb.run_all : produces one.

mtb.load_batch : reloads one from save()'s folder.

mtb.plot.bubble : the figure plot draws from long.

summary property

summary: DataFrame

One row per method: status, timing, shape, label matching and every metric.

RETURNS DESCRIPTION
DataFrame

One row per method, sorted by method. Read method, status and the metric columns (ARI, NMI ...); all columns are listed in Notes.

Examples

>>> res = mtb.load_batch("out/")
>>> res.summary[["method", "status", "ARI", "label_order_confidence"]]
>>> ok = res.summary.query("status == 'CHAIN_OK'")
>>> ok.sort_values("ARI", ascending=False)
Notes

Column reference. When nothing ran the frame is empty, with the columns below except the metrics:

method                  method id
status                  outcome; see "Status values"
run_sec                 wall-clock seconds of the method run
output_kind             embedding / graph
emb_shape               [cells, dims] of the embedding; None without one
n_tunable               number of command-line hyperparameters
label_order             label file(s) the metrics used, in order
label_order_confidence  how clearly that ordering won, 0-1
batch_source            'file_of_origin' / 'user' / None
n_batches               distinct batch values used (1 = none)
ARI, NMI, ASW, ...      one column per metric
label_order_note        why label_order_confidence is blank
caveat                  scan's caveat for the method, or ""
reason                  why a SKIPPED method did not run, or ""

A caveat of NaN: the record was saved before this column existed; n_batches and label_order still show which batches were read.

Status values.

  • CHAIN_OK - ran and scored.
  • CHAIN_OK_GRAPH_METHOD - a graph method, scored via a secondary embedding.
  • RUN_OK_NO_EMBEDDING - ran, but the method emits only a graph, so clustering metrics do not apply.
  • RUN_OK_EVAL_FAILED - the method ran and produced an embedding, but scoring it failed; see error in failures.
  • RUN_OK_NO_LABEL_MATCH - ran, but no label file matches the embedding's cell count.
  • RUN_OK - ran with evaluate=False.
  • TIMEOUT - exceeded run_all(timeout=...).
  • FAIL - the method itself errored; see error in failures.
  • SKIPPED - blocked before the run, with the reason in the reason column; in failures only when methods= named the method.

FAIL, TIMEOUT, RUN_OK_EVAL_FAILED and RUN_OK_NO_LABEL_MATCH also appear in failures.

Graph methods. scMoMaT also writes a UMAP, which is scored: CHAIN_OK_GRAPH_METHOD. Seurat_WNN writes only a neighbour graph: RUN_OK_NO_EMBEDDING, and its emb_shape is None.

Batch columns. batch_source / n_batches say which batch vector the batch metrics (ASW_batch, GC, iLISI ...) were computed against, and how many distinct values it has (1 = none):

  • 'file_of_origin' - each cell's label file (cty1.csv -> 1, cty2.csv -> 2 ...), the rule for multi-batch datasets.
  • 'user' - the vector passed as run_all(batch=) / rescore(batch=).
  • None - a single label file, so no batch structure and clustering metrics only.

Label order. label_order is which label file(s), in which order, the metrics were computed against (e.g. rna_cty.csv+atac_cty.csv). For diagonal data the embedding stacks two disjoint cell sets in a method-specific order.

Label-order confidence. label_order_confidence is (best - max(runner_up, 0)) / best over the ARI of the candidate label orders, from 0 to 1. Near 1, one order clearly fits. Below about 0.5, two orders scored alike; check that row's label order.

Optimistic bias. When more than one ordering is possible, the reported metrics are those of the ordering with the highest ARI. So they are slightly optimistic. label_order_confidence shows how far ahead the chosen order was.

Blank confidence. The column is numeric, and a blank is NaN. So > 0.5 is False for it and .isna() finds it. It is blank in three cases, named by label_order_note:

  • "single ordering" - only one ordering was possible (normal for a paired/vertical dataset with a single cty.csv).
  • "winner at chance" - the winning ordering itself scored ARI < 0.05, so the ratio would compare two noise values.
  • "not scored" - the row has no metrics.
See Also

BatchResult.failures : the rows whose status means something went wrong.

BatchResult.results : the raw records with every ordering tried.

long property

long: DataFrame

The scores as a long table, for plotting.

This is what plot and mtb.plot.bubble consume.

RETURNS DESCRIPTION
DataFrame

Columns metric, value, method, dataset, category, clustering and source. When the scores carry scored_with, so does the table. Empty, with the first seven columns, when no method produced metrics.

Examples

>>> res = mtb.load_batch("out/")
>>> mtb.plot.bubble(res.long, metrics=["ARI", "NMI"])
>>> res.long.pivot_table(index="method", columns="metric", values="value")
Notes

Source of the rows. Each record contributes the unrounded frame run_all attached (or long.csv via mtb.load_batch) when present. Otherwise the record contributes its metrics dict. Neither that dict nor the folders that mtb.data.fetch_outputs downloads carry scored_with.

See Also

BatchResult.plot : draws the bubble figure from this frame.

mtb.to_long : the wide -> long conversion used for the metrics dict.

results property

results: list

The raw per-method records: status, out_dir, metrics and the orderings tried.

Each record keeps the method's out_dir, which rescore reads.

RETURNS DESCRIPTION
list[dict]

One dict per method. Read method, status, out_dir and metrics first; the other keys are listed in Notes.

Examples

>>> res = mtb.load_batch("out/")
>>> [r["out_dir"] for r in res.results]
>>> res.results[0].get("label_order_candidates")  # None with a single ordering
Notes

Record keys. A key is present only when it applies:

method, category, dataset   what ran, and on what
modalities                  the variant's modality roles
status                      outcome (see BatchResult.summary)
out_dir                     the method's output folder
metrics                     {metric: value}, rounded to 4 places
params_used                 the hyperparameter overrides passed
run_sec, emb_shape          timing and embedding shape
labels_used                 the label file(s) behind the metrics
label_order_candidates      every label ordering tried, with its ARI
batch_source, n_batches     the batch vector the batch metrics used
batch_file                  the file in out_dir that holds a user batch
error, traceback, note      why a method failed, was skipped or was not scored
requested                   SKIPPED records: True when methods= named it
reused                      True when skip_existing reused the output
env, output_kind, n_tunable the scan row the method ran from
caveat                      that row's caveat, or ""
data_path, multibench_version, started_at   provenance of the run
data_root                   data_path as an absolute path
scripts_commit, env_flavor, hostname        the scripts, env build and computer
_long                       internal; read BatchResult.long instead

Reused outputs. A record with reused True copies scripts_commit, env_flavor and hostname from the method's earlier record in out_dir: the values of the run that made the output. Without an earlier record they are None, 'unknown' and ''.

Label-order evidence. label_order_candidates holds every ordering tried and its ARI. summary's label_order_confidence is computed from them. It is present only when more than one ordering was possible.

See Also

BatchResult.summary : the same records as a table.

failures property

failures: DataFrame

Methods that failed, timed out or could not be scored.

run_all records failures instead of raising. Check this frame.

RETURNS DESCRIPTION
DataFrame

Columns method, status, error; empty when nothing failed.

Examples

>>> res = mtb.load_batch("out/")
>>> res.failures
>>> assert res.failures.empty, res.failures.to_string()
Notes

Statuses listed.

  • FAIL and TIMEOUT.
  • SKIPPED, for a method that methods= named and that did not run.
  • RUN_OK_EVAL_FAILED - the embedding exists, scoring it failed.
  • RUN_OK_NO_LABEL_MATCH - ran, but no label file has as many cells as the output, so nothing was scored. Check the folder with mtb.inputs_for(check=True) and mtb.labels_for.

Not listed. RUN_OK_NO_EMBEDDING: those methods ran correctly and emit a graph instead of an embedding, so there is nothing for clustering metrics to score. See summary for them.

See Also

BatchResult.summary : every method, including the ones that ran but could not be scored.

plot

plot(**kw)

Bubble figure of every method that produced metrics.

Rows are methods, best first. Circle size shows the rank within a column (bigger is better). Fill compares the values in a column: the lightest is the lowest in this figure, not zero.

PARAMETERS DESCRIPTION
**kw

Keyword arguments of mtb.plot.bubble: metrics=, methods=, order=, title=, cmap=, save= ...

default {}

RETURNS DESCRIPTION
Figure

Save it with fig.savefig("out.png").

RAISES DESCRIPTION
ValueError

Nothing was scored (see failures), or an unknown name in order= / methods=.

Examples

>>> res = mtb.load_batch("out/")
>>> fig = res.plot()
>>> fig = res.plot(metrics=["ARI", "NMI", "ASW"], title="D11 vertical")
>>> fig.savefig("D11_vertical.png", dpi=200)
Notes

Keywords. metrics= sets the column order, methods= a subset, order= the row order (unlisted methods follow best-first; unknown names raise ValueError). There is no default title: pass title= when one is wanted.

Reading the figure. Size and colour are relative to the methods in this figure, so with few methods a small gap fills the whole colour scale. Check the values in summary.

See Also

mtb.plot.bubble : the underlying function and its full keyword list.

BatchResult.long : the frame handed to it.

rescore

rescore(
    *, batch=None, labels=None, metrics=None, verbose: bool = False
) -> "BatchResult"

Score the saved outputs again with new labels, batch or metrics.

No method is re-run. Each record's embedding is read back from its out_dir.

PARAMETERS DESCRIPTION
batch

Batch ids in the order of mtb.labels_for(dataset), or a barcode-indexed Series or CSV. None = the batch run_all was given, else each cell's label file.

type array - like | Series | path | None default None

labels

One cell-type label per cell; a Series or a barcode-indexed CSV is aligned by barcode. None = search the dataset's label files again.

type array - like | Series | path | None default None

metrics

Metric family ("clustering", "batch", "all") or metric codes; None = every metric the batch structure allows.

type str | list[str] | None default None

verbose

Print one line per method, and one before each label-order ranking.

type bool default False

RETURNS DESCRIPTION
BatchResult

A new result; this one is untouched. Call .save(out_dir) on it to persist it.

RAISES DESCRIPTION
ValueError

labels or batch holds ids that are not cells of the dataset.

WARNS DESCRIPTION
UserWarning

A Series or barcode-indexed CSV cannot be aligned and is matched by position.

Examples

>>> res = mtb.load_batch("out/")
>>> res.rescore(metrics=["ARI", "NMI"]).summary
>>> new = res.rescore(batch="data/D11/donor.csv")
>>> new.summary[["method", "batch_source", "iLISI"]]
>>> res.rescore(labels=my_labels).save("out/rescored")
Notes

Arguments. A labels array, list or plain CSV follows the embedding rows. Without labels=, a batch array follows the order of mtb.labels_for(dataset). With a labels array in embedding row order, the batch array follows the embedding rows too. A CSV path is read like a label file. The batch is recorded as batch_source='user'.

Aligned by barcode. A Series or one-column DataFrame with a non-default index, or a CSV whose first column holds barcodes, is aligned to the dataset's cells as in run_all(batch=), with the same errors and warnings. Aligned labels go into each record's label-order search, so label_order names the file order chosen. Labels matched by position read (user labels).

Label order. With labels=None and several label files, rescore ranks the file orders by ARI, which needs one Leiden sweep. It keeps the stored order and skips the sweep when metrics= has no ARI, NMI or iF1.

Labels without batch. Given labels and no batch, the batch that run_all(batch=) saved is reused. Without a saved batch, every cell is in one batch (batch_source None, n_batches 1), so only clustering metrics are computed. Pass batch as well to get the batch metrics.

Record status. A method that emits no embedding (graph-only) is marked RUN_OK_NO_EMBEDDING with a note. SKIPPED, FAIL and TIMEOUT records are kept as they are. A record whose output file is gone becomes RUN_OK_EVAL_FAILED, with the reason in error. So does a record whose new scoring fails, for example on a batch of the wrong length (batch has N entries, embedding has M cells).

Other hosts. When the dataset folder has moved, pass the folder that now holds it as mtb.load_batch(data_path=). When the dataset folder is not found, labels=None gives RUN_OK_NO_LABEL_MATCH with a note. A Series or CSV is then matched by position, with a warning, and the saved batch is not used.

Persisting. mtb.load_batch keeps returning the original result until the new one is saved.

See Also

mtb.evaluate : the scoring function applied per record.

BatchResult.save : persist the re-scored result.

save

save(out_dir=None) -> 'Path'

Write this result to disk so it outlives the process.

PARAMETERS DESCRIPTION
out_dir

Target folder, created if missing; None = the result's own out_dir, else the current directory.

type path | None default None

RETURNS DESCRIPTION
Path

The folder written.

RAISES DESCRIPTION
ValueError

The folder holds a saved result of another dataset or category.

Examples

>>> res = mtb.load_batch("out/")
>>> res.rescore(metrics="clustering").save("out/clustering_only")
>>> mtb.load_batch("out/clustering_only").summary
Notes

Files written.

  • summary.csv - the summary frame.
  • long.csv - the long frame, only when some method produced metrics.
  • failures.csv - the failures frame.
  • batch_result.json - dataset, category and the per-method records.
  • batch_<hash>.csv - a batch vector the records were scored with.

Reload it with mtb.load_batch.

Saving into a folder that has a result. When the folder already holds batch_result.json for the same dataset and category, the records are merged. A method in this result replaces its earlier record, unless it is SKIPPED and the earlier one is not; the other earlier records are kept, and every file in the list above is rewritten from the merged set. This result object is not changed.

A line # Merged with 1 earlier record in <folder> (StabMap). names the kept methods.

Jobs in parallel. Jobs running at the same time should use one out_dir each; combine them with multibench plot --input dir1 --input dir2.

Blank confidence on disk. In summary.csv the label_order_note column says why label_order_confidence is empty on a row: "single ordering", "winner at chance" or "not scored".

See Also

mtb.load_batch : reads the folder back.