Skip to content

Discover

Find methods, see what each one needs and cite them. Every call works offline.

Name Summary
mtb.list_methods Return the method ids, optionally restricted to one category.
mtb.find_methods Return the method ids that match every filter you pass.
mtb.method_info Return everything known about one method as a flat dict.
mtb.params_for Return the hyperparameters of one method variant.
mtb.describe_layout Return the directory layout the package expects for your own dataset.
mtb.cite Return citation text for the benchmark and, optionally, the methods you ran.
mtb.list_tasks List the task names that methods declare, sorted.
mtb.list_categories Return the valid category values with a plain-language description of each.

mtb.list_methods

list_methods(category: str | None = None) -> list[str]

Return the method ids, optionally restricted to one category.

PARAMETERS DESCRIPTION
category

Integration category: vertical, diagonal, mosaic or cross; None = every method.

type str | None default None

RETURNS DESCRIPTION
list[str]

Method ids in the package's order.

RAISES DESCRIPTION
ValueError

Unknown category; the message lists the valid ones.

TypeError

A keyword other than category; see Other keywords in Notes.

Examples

>>> import multibench as mtb
>>> mtb.list_methods()                 # every method id
>>> mtb.list_methods("vertical")       # ids with a vertical variant
Notes

Category membership. A method is listed under a category when it has a variant for that category - the same set scan, run_all and find_methods(category=) dispatch.

Other keywords. A find_methods filter such as task= or runnable= raises a TypeError that names the call to use instead: find_methods(category, task=...). Any other keyword raises Python's own unexpected keyword argument message.

See Also

mtb.find_methods : filter by task, modalities, labels, ATAC representation and more.

mtb.list_categories : the four category tokens with a description of each.

mtb.find_methods

find_methods(
    category: str | None = None,
    *,
    task: str | None = None,
    needs_labels: bool | None = None,
    atac: str | None = None,
    modalities: list[str] | set[str] | None = None,
    runnable: bool | None = None,
    tunable: bool | None = None,
) -> list[str]

Return the method ids that match every filter you pass.

A method matches when one of its variants meets category, modalities, needs_labels and atac together.

PARAMETERS DESCRIPTION
category

Integration category: vertical, diagonal, mosaic or cross; None = any.

type str | None default None

task

A task from mtb.list_tasks(), e.g. "clustering"; None = any.

type str | None default None

needs_labels

True = a matching variant needs cell-type labels; False = one runs without them; None = no filter.

type bool | None default None

atac

ATAC representation the script expects: "peak" or "gene_activity"; None = no filter.

type str | None default None

modalities

Modalities a variant must consume, e.g. ["rna", "adt"]; "atac_peak" / "atac_gas" also select by ATAC representation. None = no filter.

type list[str] | set[str] | None default None

runnable

True = methods with a variant the package runs (not a check of your host: see mtb.scan); False = declared stubs; None = both.

type bool | None default None

tunable

True = methods with command-line hyperparameters that run(params=...) can set; False = the rest (settings fixed in the script); None = both.

type bool | None default None

RETURNS DESCRIPTION
list[str]

Method ids in the package's order.

RAISES DESCRIPTION
ValueError

Unknown category, task, atac or modality token, or a token that contradicts atac.

TypeError

modalities is a bare string, not a list.

Examples

>>> import multibench as mtb
>>> mtb.find_methods("vertical", modalities=["rna", "adt"])
>>> mtb.find_methods("vertical", modalities=["rna", "adt"], needs_labels=False)
>>> mtb.find_methods("diagonal", modalities=["rna", "atac_peak"])
>>> mtb.find_methods(atac="gene_activity")
>>> mtb.find_methods(tunable=True)
Notes

Per-variant matching. task, runnable and tunable are method-level; the other four filters hold per variant. Two consequences:

  • find_methods('vertical', modalities=['rna', 'adt'], needs_labels=False) keeps scMoMaT: its vertical rna+adt variant takes no labels; only its mosaic variant does.
  • find_methods('vertical', modalities=['rna', 'atac']) drops Multigrate: rna+atac exists only as a mosaic variant, so inputs_for(..., 'vertical', modalities=['rna', 'atac']) would raise.

Modality tokens. A method matches when one variant reads at least the named modalities. mtb.scan and mtb.run_all use a stricter rule: a row's modalities must be exactly the named combination.

A base token matches every role of its type: atac matches the atac, atac_gas and atac_peak roles, rna matches rna1, and protein is adt.

A representation token selects by what the method reads. atac_peak (also peak) is atac plus atac='peak'; atac_gas (also gas, gene_activity) is atac plus atac='gene_activity'. Both tokens together select the variants that read both files (MultiMAP, Seurat_v3). A representation that contradicts atac= raises ValueError.

Directory-fed methods. scBridge is fed a directory and is judged by the bare filenames its variant names: rna.h5 / atac_gas.h5 make it an rna+atac method.

ATAC representation. atac is what the upstream script reads, which is not always what its role name suggests: moETM, scMM and iPOLNG take the role atac_gas but read peaks. MultiMAP and Seurat_v3 read peaks and gene activity; they are listed under peak.

Only a variant with an ATAC input satisfies atac: Multigrate reads peaks only in its mosaic rna+atac variant, so find_methods('vertical', atac='peak') omits it. Accepted spellings: peak / peaks and gene_activity / gene-activity / gas, in any case.

Labels. needs_labels here is per variant. method_info(m)['needs_labels'] is the method-level flag (any variant needs labels); the per-variant answer is method_info(m)['supports'][i]['needs_labels'].

Stubs. A declared stub (a method with no variant) is dropped by any filter that asks something of a variant: category, modalities, atac, needs_labels=True or tunable=True.

See Also

mtb.list_methods : the same ids filtered by category only.

mtb.list_tasks : the vocabulary of task=.

mtb.method_info : the per-variant supports table these filters are read from.

mtb.scan : which of these methods can actually run on a given dataset.

mtb.method_info

method_info(method: str, *, verbose: bool = False) -> dict

Return everything known about one method as a flat dict.

PARAMETERS DESCRIPTION
method

Method id, e.g. "Matilda"; see mtb.list_methods().

type str

verbose

True adds notes_long, the long upstream-audit notes.

type bool default False

RETURNS DESCRIPTION
dict

One flat record. Start with supports, params, runtime and needs_labels. Notes describes every key.

RAISES DESCRIPTION
KeyError

Unknown method id. The message suggests a close match.

Examples

>>> import multibench as mtb
>>> info = mtb.method_info("Matilda")
>>> # one entry per variant: category, modalities, labels ...
>>> info["supports"]
>>> info["runtime"]["tier"], info["runtime"]["worst_sec"]
>>> info["params"]["vertical:rna+adt"]["tunable"]
Notes

Key reference. The dict merges the method definition, the provenance record (repository, version, paper) and the observed runtime.

  • id - the method id.
  • language - 'python' or 'R'.
  • categories / tasks - the categories it has a variant for and the tasks it serves.
  • env - the conda env run executes it in (a shared group env or the method's own).
  • atac - the ATAC representation the script expects ('peak' / 'gene_activity'), or None.
  • needs_labels - method-level label flag; see Labels below.
  • status - see Status.
  • setup_hint - free-text setup advice, or '' when there is none.
  • variants - the distinct upstream entrypoints, in order.
  • driver - the package-side wrapper actually executed, or None when the upstream script runs directly.
  • scripts_url - the method's tools_scripts folder in the scMultiBench repository, percent-encoded.
  • repo_url / version - the upstream repository and the version the benchmark ran.
  • reference - {doi, title, authors, journal, year} or None; mtb.cite formats it.
  • notes - a short summary of the method.
  • supports - the category and modality combinations the method runs, one entry per variant: category, modalities, output_kind, n_tunable, needs_labels, labels (the label roles the variant reads, e.g. ['cty'] / ['rna_cty'] / []) and reference_batch (the batch a variant uses as its fixed reference, e.g. 3 for StabMap in cross; None elsewhere).
  • params - what run(params=...) can change, keyed per variant as 'category:mods', each with defaults, tunable and effective (see mtb.params_for).
  • fixed_in_script / upstream_knobs / upstream_url - what the script pins and what its library documents (see mtb.params_for).
  • runtime - observed cost, to size a sweep; see Runtime.
  • gpu / cpu_params / requires_gpu / gpu_evidence - see GPU and CPU.
  • notes_long (verbose=True only) - the raw upstream-knob audit prose, None for methods outside the audit.
  • Not in this dict: the paper-only catalog columns (deep_learning, output); read them from mtb.catalog.methods().

Runtime. runtime is {"tier", "worst_sec", "observed", "host", "note"}, what this method has been observed to cost:

  • tier - fast (<5 min), medium (5-30 min), slow (30 min-2 h), very_slow (>2 h) or unknown (never measured: worst_sec is None, observed empty).
  • worst_sec - the slowest observation, in seconds.
  • observed - one {dataset, cells, sec, source} per measurement; cells is None when not recorded; source is manual, summary_csv (the shipped re-run sweeps) or recorded (the recorded end-to-end runs).
  • host / note - 'gpu' and a sentence saying the times come from the GPU benchmark host; for a method that uses a GPU it adds that a CPU-only host takes longer.

Use them to set run_all(timeout=...).

Labels. needs_labels is the method-level flag: True when any variant takes a cell-type-label (cty) role as a required input. It is not per category: scMoMaT is True because its mosaic variant takes cty1..3, while its vertical and cross variants take no labels. For the per-variant answer read supports[i]['needs_labels']; find_methods(needs_labels=...) filters per variant.

Status. status says how the method was checked. 'verified' means the command template was cross-checked against the upstream entrypoint and the method was executed end to end on a reference dataset; 'declared' = available but not run end to end. A method listed but not yet runnable raises KeyError.

GPU and CPU. gpu, cpu_params, requires_gpu and gpu_evidence are the GPU/CPU contract of the upstream script, read from its source:

  • gpu - how the method uses an NVIDIA GPU: 'required' (see requires_gpu), 'used when present' (it runs on the CPU otherwise), 'not used', or 'unknown' (not checked yet).
  • cpu_params - the command-line values that turn CUDA off in a script that has it on by default ({} for most methods): {'use_cuda': ''} for scJoint (its argparse --use_cuda is type=bool, so only the empty string is false), {'device': 'cpu'} for scMDC. On a host without an NVIDIA GPU (mtb.env.host_has_gpu() False) run merges them into params unless the caller set the key.
  • requires_gpu - True for a script that calls CUDA unconditionally (no flag, no torch.cuda.is_available() fallback). On a GPU-less host run refuses such a method with OSError before launching and scan reports it env_ok=False.
  • gpu_evidence - the file:line of that CUDA call (one per script, joined by ', '), else None.

None of the three says anything about the CPU archive of the method's env (mtb.env.install(..., flavor=...)).

See Also

mtb.params_for : the hyperparameters of one variant, with upstream defaults.

mtb.find_methods : filter methods by category, modalities, labels, ATAC representation.

mtb.cite : the paper to cite for a method.

mtb.env.recipe : the hand-written environment recipe of a method.

mtb.params_for

params_for(
    method: str,
    category: str | None = None,
    modalities: list[str] | set[str] | None = None,
    *,
    dataset: str | None = None,
    data_path: Path | str | None = None,
) -> dict

Return the hyperparameters of one method variant.

Call it before run(params=...), run_all(params=...) or sweep to see what a method accepts and what it runs with when you pass nothing.

PARAMETERS DESCRIPTION
method

Method id, e.g. "Matilda"; see mtb.list_methods().

type str

category

Integration category of the variant: vertical, diagonal, mosaic or cross; None = infer it from the other arguments.

type str | None default None

modalities

Modality tokens of the variant, e.g. ["rna", "adt"]; None = infer it from the other arguments.

type list[str] | set[str] | None default None

dataset

Dataset folder that settles an ambiguous selection: the one variant whose input files it holds.

type str | None default None

data_path

Data root that holds the dataset folders; None = mtb.config.DEFAULT.data_path.

type Path | str | None default None

RETURNS DESCRIPTION
dict

method, variant and the hyperparameters. Read tunable (what the script accepts) and effective (what a run uses); all keys are listed in Notes.

RAISES DESCRIPTION
KeyError

Unknown method, a declared stub, or no variant for category / modalities.

AmbiguousVariantError

Several variants fit; the message spells out the call that selects one.

ValueError

Unknown category or modality token.

Examples

>>> import multibench as mtb
>>> p = mtb.params_for("Matilda", "vertical", ["rna", "adt"])
>>> # upstream default vs what a run uses
>>> p["tunable"]["device"]["default"], p["effective"]["device"]
>>> mtb.params_for("Matilda", dataset="D11")  # the folder picks the variant
>>> mtb.params_for("scBridge", "diagonal")  # a data_dir variant: no modalities
Notes

Key reference.

  • method / variant - the method id and the selected variant as 'category:mods' (mods is - for a data_dir variant, e.g. 'diagonal:-').
  • defaults - parameters the package emits on every run. Override them with run(..., params={...}); the override is merged over these.
  • tunable - the parameters the upstream script accepts on its command line, as {name: {"default": ..., "type": ...}}. The default here is the upstream argparse default, not necessarily what a wrapper run uses.
  • effective - tunable defaults overlaid with defaults: the value each knob really takes when you pass no params.
  • fixed_in_script - the values the script pins, each with the file:line that pins it.
  • upstream_knobs - what the wrapped library documents (with its own defaults), unreachable without editing the script.
  • upstream_url - the upstream source or docs page those knobs were read from, or None.

Methods with nothing to tune. An empty tunable means the upstream script exposes no hyperparameters on its command line. The package runs each script unchanged, so such a method cannot be tuned through the wrapper. Its settings are reported under fixed_in_script and upstream_knobs (both empty for methods outside the upstream audit).

Variant selection. category and modalities select the variant exactly like run. Either may be omitted when the rest leaves one variant. A data_dir variant (scBridge's) has no modality tokens: leave modalities out and select it by category, or by nothing when it is the method's only variant.

Dataset tie-break. When the selection is still ambiguous, dataset picks the one variant whose input files are all present in <data_path>/<dataset>: params_for('Matilda', dataset='D11') is the rna+adt variant. A folder that settles nothing changes nothing, and the ambiguity error is raised as usual.

Modality spellings. protein is accepted for adt, and atac for either ATAC representation role (atac_gas / atac_peak).

Ambiguity error. mtb.AmbiguousVariantError derives from both ValueError and KeyError. Its message spells out the selecting call, e.g. params_for('Matilda', 'vertical', ['rna', 'adt']).

See Also

mtb.method_info : the same params block for every variant at once.

mtb.sweep : run one method over a range of one of these hyperparameters.

mtb.AmbiguousVariantError : raised when several variants fit the selection.

mtb.describe_layout

describe_layout(category: str | None = None) -> str

Return the directory layout the package expects for your own dataset.

Start here when bringing your own data, then check the folder with mtb.scan.

PARAMETERS DESCRIPTION
category

Integration category to describe; None = an overview of all four and the full file table.

type str | None default None

RETURNS DESCRIPTION
str

The layout description, ready to print.

RAISES DESCRIPTION
ValueError

Unknown category.

Examples

>>> import multibench as mtb
>>> print(mtb.describe_layout("vertical"))  # CITE-seq or multiome
>>> print(mtb.describe_layout("mosaic"))    # batch patterns methods accept
>>> print(mtb.describe_layout())            # every category
Notes

What the text covers. For one category: the files its methods read, one name per file, as mtb.io.export_dataset writes them. It also gives the ATAC representation each method needs, the .h5 and label formats, and the install command. For mosaic and cross it lists each batch pattern the methods accept. multibench layout prints the same text, with multibench commands in place of the Python calls.

One rule for ATAC files. Vertical reads atac.h5; method_info(m)["atac"] says whether it must hold peaks or gene activity. Diagonal reads atac_peak.h5 (peaks) and atac_gas.h5 (gene activity). Mosaic reads atac<i>.h5 (peaks). peak.h5, and atac.h5 for gene activity, are accepted as older names.

Several batches. Mosaic and cross use one numbered file per batch in the same folder, not sub-folders and not one concatenated matrix. The file number is the batch; there is no batch column:

<data_path>/COREBATCH/
    rna1.h5   adt1.h5   cty1.csv     # batch 1
    rna2.h5   adt2.h5   cty2.csv     # batch 2
    rna3.h5   adt3.h5   cty3.csv     # batch 3

Source of the lists. The methods per ATAC representation and the batch patterns are read from the package's method list at call time, so they agree with method_info and mtb.find_methods.

See Also

mtb.list_categories : the four categories with a description of each.

mtb.scan : checks a laid-out folder (files_ok / files_reason per method).

mtb.io.export_dataset : writes a whole dataset in this layout from an AnnData.

mtb.cite

cite(*methods, fmt: str = 'text') -> str

Return citation text for the benchmark and, optionally, the methods you ran.

PARAMETERS DESCRIPTION
*methods

Method ids, one per argument or as one list; "all" = every method; none = the benchmark only.

type str | list[str] default ()

fmt

"text" (one line per entry) or "bibtex" (one @article per entry).

type str default 'text'

RETURNS DESCRIPTION
str

The benchmark entry first, then one entry per method in the order given.

RAISES DESCRIPTION
ValueError

fmt is neither "text" nor "bibtex"; the message lists both.

KeyError

Unknown method id; the message suggests a close match, if any.

TypeError

Several ids are given and one is not a string.

Examples

>>> import multibench as mtb
>>> print(mtb.cite())                                  # the benchmark only
>>> print(mtb.cite("Matilda", "MOFA2"))
>>> res = mtb.load_batch("out/")
>>> ran = res.summary.query("status not in ['SKIPPED', 'FAIL', 'TIMEOUT']")
>>> print(mtb.cite(ran.method, fmt="bibtex"))      # methods that finished
Notes

Two spellings. Both forms return the same text, and cite(None) is cite():

  • cite('Matilda', 'MOFA2') - one id per argument, like the CLI multibench cite Matilda MOFA2.
  • cite(['Matilda', 'MOFA2']) - one list or tuple.

Formats.

  • "text" - one Authors. Title. Journal (year). https://doi.org/... line per entry.
  • "bibtex" - one @article{<id>_<year>, ...} per entry, separated by a blank line; the benchmark's key is scMultiBench_<year>.

Methods without a reference. A method without a verified reference is emitted as a % <id>: no verified reference; see <repo_url> comment (bibtex) or the same line without % (text).

See Also

mtb.method_info : carries the same reference, repo_url and version per method.

mtb.list_tasks

list_tasks() -> list[str]

List the task names that methods declare, sorted.

RETURNS DESCRIPTION
list[str]

Sorted task names, the values mtb.find_methods(task=...) accepts.

Examples

>>> import multibench as mtb
>>> mtb.list_tasks()  # ['batch', 'clustering', 'dimension_reduction']
>>> mtb.find_methods(task="dimension_reduction")
See Also

mtb.find_methods : filter methods by one of these tasks.

mtb.list_categories

list_categories() -> dict

Return the valid category values with a plain-language description of each.

RETURNS DESCRIPTION
dict

{category: description} for vertical, diagonal, mosaic and cross: the values mtb.run_all requires as category.

Examples

>>> import multibench as mtb
>>> mtb.list_categories()["vertical"]
'Several modalities measured in the same cells ...'
See Also

mtb.describe_layout : the file layout each category expects.