Discover¶
Find methods, see what each one needs and cite them. Every call works offline.
| Name | Summary |
|---|---|
mtb.list_methods |
Return the method ids, optionally restricted to one category. |
mtb.find_methods |
Return the method ids that match every filter you pass. |
mtb.method_info |
Return everything known about one method as a flat dict. |
mtb.params_for |
Return the hyperparameters of one method variant. |
mtb.describe_layout |
Return the directory layout the package expects for your own dataset. |
mtb.cite |
Return citation text for the benchmark and, optionally, the methods you ran. |
mtb.list_tasks |
List the task names that methods declare, sorted. |
mtb.list_categories |
Return the valid category values with a plain-language description of each. |
mtb.list_methods
¶
Return the method ids, optionally restricted to one category.
| PARAMETERS | DESCRIPTION |
|---|---|
category
|
Integration category:
type
|
| RETURNS | DESCRIPTION |
|---|---|
list[str]
|
Method ids in the package's order. |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
Unknown |
TypeError
|
A keyword other than |
Examples
>>> import multibench as mtb
>>> mtb.list_methods() # every method id
>>> mtb.list_methods("vertical") # ids with a vertical variant
Notes
Category membership. A method is listed under a category when it has
a variant for that category - the same set scan, run_all
and find_methods(category=) dispatch.
Other keywords. A find_methods filter such as task= or
runnable= raises a TypeError that names the call to use
instead: find_methods(category, task=...). Any other keyword raises
Python's own unexpected keyword argument message.
See Also
mtb.find_methods : filter by task, modalities, labels, ATAC representation and more.
mtb.list_categories : the four category tokens with a description of each.
mtb.find_methods
¶
find_methods(
category: str | None = None,
*,
task: str | None = None,
needs_labels: bool | None = None,
atac: str | None = None,
modalities: list[str] | set[str] | None = None,
runnable: bool | None = None,
tunable: bool | None = None,
) -> list[str]
Return the method ids that match every filter you pass.
A method matches when one of its variants meets category,
modalities, needs_labels and atac together.
| PARAMETERS | DESCRIPTION |
|---|---|
category
|
Integration category:
type
|
task
|
A task from
type
|
needs_labels
|
type
|
atac
|
ATAC representation the script expects:
type
|
modalities
|
Modalities a variant must consume, e.g.
type
|
runnable
|
type
|
tunable
|
type
|
| RETURNS | DESCRIPTION |
|---|---|
list[str]
|
Method ids in the package's order. |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
Unknown |
TypeError
|
|
Examples
>>> import multibench as mtb
>>> mtb.find_methods("vertical", modalities=["rna", "adt"])
>>> mtb.find_methods("vertical", modalities=["rna", "adt"], needs_labels=False)
>>> mtb.find_methods("diagonal", modalities=["rna", "atac_peak"])
>>> mtb.find_methods(atac="gene_activity")
>>> mtb.find_methods(tunable=True)
Notes
Per-variant matching. task, runnable and tunable are
method-level; the other four filters hold per variant. Two consequences:
find_methods('vertical', modalities=['rna', 'adt'], needs_labels=False)keeps scMoMaT: its vertical rna+adt variant takes no labels; only its mosaic variant does.find_methods('vertical', modalities=['rna', 'atac'])drops Multigrate: rna+atac exists only as a mosaic variant, soinputs_for(..., 'vertical', modalities=['rna', 'atac'])would raise.
Modality tokens. A method matches when one variant reads at least
the named modalities. mtb.scan and mtb.run_all use a stricter
rule: a row's modalities must be exactly the named combination.
A base token matches every role of its type: atac matches the
atac, atac_gas and atac_peak roles, rna matches rna1,
and protein is adt.
A representation token selects by what the method reads. atac_peak
(also peak) is atac plus atac='peak'; atac_gas (also
gas, gene_activity) is atac plus atac='gene_activity'.
Both tokens together select the variants that read both files
(MultiMAP, Seurat_v3). A representation that contradicts atac=
raises ValueError.
Directory-fed methods. scBridge is fed a directory and is judged by
the bare filenames its variant names: rna.h5 / atac_gas.h5 make it
an rna+atac method.
ATAC representation. atac is what the upstream script reads,
which is not always what its role name suggests: moETM, scMM and iPOLNG
take the role atac_gas but read peaks. MultiMAP and Seurat_v3 read
peaks and gene activity; they are listed under peak.
Only a variant with an ATAC input satisfies atac: Multigrate reads
peaks only in its mosaic rna+atac variant, so
find_methods('vertical', atac='peak') omits it. Accepted spellings:
peak / peaks and gene_activity / gene-activity / gas,
in any case.
Labels. needs_labels here is per variant.
method_info(m)['needs_labels'] is the method-level flag (any variant
needs labels); the per-variant answer is
method_info(m)['supports'][i]['needs_labels'].
Stubs. A declared stub (a method with no variant) is dropped by any
filter that asks something of a variant: category, modalities,
atac, needs_labels=True or tunable=True.
See Also
mtb.list_methods : the same ids filtered by category only.
mtb.list_tasks : the vocabulary of task=.
mtb.method_info : the per-variant supports table these filters are read from.
mtb.scan : which of these methods can actually run on a given dataset.
mtb.method_info
¶
Return everything known about one method as a flat dict.
| PARAMETERS | DESCRIPTION |
|---|---|
method
|
Method id, e.g.
type
|
verbose
|
type
|
| RETURNS | DESCRIPTION |
|---|---|
dict
|
One flat record. Start with |
| RAISES | DESCRIPTION |
|---|---|
KeyError
|
Unknown method id. The message suggests a close match. |
Examples
>>> import multibench as mtb
>>> info = mtb.method_info("Matilda")
>>> # one entry per variant: category, modalities, labels ...
>>> info["supports"]
>>> info["runtime"]["tier"], info["runtime"]["worst_sec"]
>>> info["params"]["vertical:rna+adt"]["tunable"]
Notes
Key reference. The dict merges the method definition, the provenance record (repository, version, paper) and the observed runtime.
id- the method id.language-'python'or'R'.categories/tasks- the categories it has a variant for and the tasks it serves.env- the conda envrunexecutes it in (a shared group env or the method's own).atac- the ATAC representation the script expects ('peak'/'gene_activity'), orNone.needs_labels- method-level label flag; see Labels below.status- see Status.setup_hint- free-text setup advice, or''when there is none.variants- the distinct upstream entrypoints, in order.driver- the package-side wrapper actually executed, orNonewhen the upstream script runs directly.scripts_url- the method'stools_scriptsfolder in the scMultiBench repository, percent-encoded.repo_url/version- the upstream repository and the version the benchmark ran.reference-{doi, title, authors, journal, year}orNone;mtb.citeformats it.notes- a short summary of the method.supports- the category and modality combinations the method runs, one entry per variant:category,modalities,output_kind,n_tunable,needs_labels,labels(the label roles the variant reads, e.g.['cty']/['rna_cty']/[]) andreference_batch(the batch a variant uses as its fixed reference, e.g. 3 for StabMap in cross;Noneelsewhere).params- whatrun(params=...)can change, keyed per variant as'category:mods', each withdefaults,tunableandeffective(seemtb.params_for).fixed_in_script/upstream_knobs/upstream_url- what the script pins and what its library documents (seemtb.params_for).runtime- observed cost, to size a sweep; see Runtime.gpu/cpu_params/requires_gpu/gpu_evidence- see GPU and CPU.notes_long(verbose=Trueonly) - the raw upstream-knob audit prose,Nonefor methods outside the audit.- Not in this dict: the paper-only catalog columns (
deep_learning,output); read them frommtb.catalog.methods().
Runtime. runtime is {"tier", "worst_sec", "observed", "host",
"note"}, what this method has been observed to cost:
tier-fast(<5 min),medium(5-30 min),slow(30 min-2 h),very_slow(>2 h) orunknown(never measured:worst_secisNone,observedempty).worst_sec- the slowest observation, in seconds.observed- one{dataset, cells, sec, source}per measurement;cellsisNonewhen not recorded;sourceismanual,summary_csv(the shipped re-run sweeps) orrecorded(the recorded end-to-end runs).host/note-'gpu'and a sentence saying the times come from the GPU benchmark host; for a method that uses a GPU it adds that a CPU-only host takes longer.
Use them to set run_all(timeout=...).
Labels. needs_labels is the method-level flag: True when
any variant takes a cell-type-label (cty) role as a required input.
It is not per category: scMoMaT is True because its mosaic variant takes
cty1..3, while its vertical and cross variants take no labels. For
the per-variant answer read supports[i]['needs_labels'];
find_methods(needs_labels=...) filters per variant.
Status. status says how the method was checked. 'verified'
means the command template was cross-checked against the upstream entrypoint
and the method was executed end to end on a reference dataset;
'declared' = available but not run end to end. A method listed but
not yet runnable raises KeyError.
GPU and CPU. gpu, cpu_params, requires_gpu and
gpu_evidence are the GPU/CPU contract of the upstream script, read
from its source:
gpu- how the method uses an NVIDIA GPU:'required'(seerequires_gpu),'used when present'(it runs on the CPU otherwise),'not used', or'unknown'(not checked yet).cpu_params- the command-line values that turn CUDA off in a script that has it on by default ({}for most methods):{'use_cuda': ''}for scJoint (its argparse--use_cudaistype=bool, so only the empty string is false),{'device': 'cpu'}for scMDC. On a host without an NVIDIA GPU (mtb.env.host_has_gpu()False)runmerges them intoparamsunless the caller set the key.requires_gpu-Truefor a script that calls CUDA unconditionally (no flag, notorch.cuda.is_available()fallback). On a GPU-less hostrunrefuses such a method withOSErrorbefore launching andscanreports itenv_ok=False.gpu_evidence- thefile:lineof that CUDA call (one per script, joined by', '), elseNone.
None of the three says anything about the CPU archive of the method's env
(mtb.env.install(..., flavor=...)).
See Also
mtb.params_for : the hyperparameters of one variant, with upstream defaults.
mtb.find_methods : filter methods by category, modalities, labels, ATAC representation.
mtb.cite : the paper to cite for a method.
mtb.env.recipe : the hand-written environment recipe of a method.
mtb.params_for
¶
params_for(
method: str,
category: str | None = None,
modalities: list[str] | set[str] | None = None,
*,
dataset: str | None = None,
data_path: Path | str | None = None,
) -> dict
Return the hyperparameters of one method variant.
Call it before run(params=...), run_all(params=...) or sweep to
see what a method accepts and what it runs with when you pass nothing.
| PARAMETERS | DESCRIPTION |
|---|---|
method
|
Method id, e.g.
type
|
category
|
Integration category of the variant:
type
|
modalities
|
Modality tokens of the variant, e.g.
type
|
dataset
|
Dataset folder that settles an ambiguous selection: the one variant whose input files it holds.
type
|
data_path
|
Data root that holds the dataset folders;
type
|
| RETURNS | DESCRIPTION |
|---|---|
dict
|
|
| RAISES | DESCRIPTION |
|---|---|
KeyError
|
Unknown method, a declared stub, or no variant for |
AmbiguousVariantError
|
Several variants fit; the message spells out the call that selects one. |
ValueError
|
Unknown |
Examples
>>> import multibench as mtb
>>> p = mtb.params_for("Matilda", "vertical", ["rna", "adt"])
>>> # upstream default vs what a run uses
>>> p["tunable"]["device"]["default"], p["effective"]["device"]
>>> mtb.params_for("Matilda", dataset="D11") # the folder picks the variant
>>> mtb.params_for("scBridge", "diagonal") # a data_dir variant: no modalities
Notes
Key reference.
method/variant- the method id and the selected variant as'category:mods'(modsis-for adata_dirvariant, e.g.'diagonal:-').defaults- parameters the package emits on every run. Override them withrun(..., params={...}); the override is merged over these.tunable- the parameters the upstream script accepts on its command line, as{name: {"default": ..., "type": ...}}. Thedefaulthere is the upstream argparse default, not necessarily what a wrapper run uses.effective-tunabledefaults overlaid withdefaults: the value each knob really takes when you pass noparams.fixed_in_script- the values the script pins, each with thefile:linethat pins it.upstream_knobs- what the wrapped library documents (with its own defaults), unreachable without editing the script.upstream_url- the upstream source or docs page those knobs were read from, orNone.
Methods with nothing to tune. An empty tunable means the upstream
script exposes no hyperparameters on its command line. The package runs
each script unchanged, so such a method cannot be tuned through the
wrapper. Its settings are reported under fixed_in_script and
upstream_knobs (both empty for methods outside the upstream audit).
Variant selection. category and modalities select the variant
exactly like run. Either may be omitted when the rest leaves one
variant. A data_dir variant (scBridge's) has no modality
tokens: leave modalities out and select it by category, or by
nothing when it is the method's only variant.
Dataset tie-break. When the selection is still ambiguous, dataset
picks the one variant whose input files are all present in
<data_path>/<dataset>: params_for('Matilda', dataset='D11') is the
rna+adt variant. A folder that settles nothing changes nothing, and the
ambiguity error is raised as usual.
Modality spellings. protein is accepted for adt, and atac
for either ATAC representation role (atac_gas / atac_peak).
Ambiguity error. mtb.AmbiguousVariantError derives from both
ValueError and KeyError. Its message spells out the selecting
call, e.g. params_for('Matilda', 'vertical', ['rna', 'adt']).
See Also
mtb.method_info : the same params block for every variant at once.
mtb.sweep : run one method over a range of one of these hyperparameters.
mtb.AmbiguousVariantError : raised when several variants fit the selection.
mtb.describe_layout
¶
Return the directory layout the package expects for your own dataset.
Start here when bringing your own data, then check the folder with
mtb.scan.
| PARAMETERS | DESCRIPTION |
|---|---|
category
|
Integration category to describe;
type
|
| RETURNS | DESCRIPTION |
|---|---|
str
|
The layout description, ready to |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
Unknown |
Examples
>>> import multibench as mtb
>>> print(mtb.describe_layout("vertical")) # CITE-seq or multiome
>>> print(mtb.describe_layout("mosaic")) # batch patterns methods accept
>>> print(mtb.describe_layout()) # every category
Notes
What the text covers. For one category: the files its methods read,
one name per file, as mtb.io.export_dataset writes them. It also
gives the ATAC representation each method needs, the .h5 and label
formats, and the install command. For mosaic and cross it lists each
batch pattern the methods accept. multibench layout prints the same
text, with multibench commands in place of the Python calls.
One rule for ATAC files. Vertical reads atac.h5;
method_info(m)["atac"] says whether it must hold peaks or gene
activity. Diagonal reads atac_peak.h5 (peaks) and atac_gas.h5
(gene activity). Mosaic reads atac<i>.h5 (peaks). peak.h5, and
atac.h5 for gene activity, are accepted as older names.
Several batches. Mosaic and cross use one numbered file per batch in the same folder, not sub-folders and not one concatenated matrix. The file number is the batch; there is no batch column:
<data_path>/COREBATCH/
rna1.h5 adt1.h5 cty1.csv # batch 1
rna2.h5 adt2.h5 cty2.csv # batch 2
rna3.h5 adt3.h5 cty3.csv # batch 3
Source of the lists. The methods per ATAC representation and the
batch patterns are read from the package's method list at call time, so they
agree with method_info and mtb.find_methods.
See Also
mtb.list_categories : the four categories with a description of each.
mtb.scan : checks a laid-out folder (files_ok / files_reason per method).
mtb.io.export_dataset : writes a whole dataset in this layout from an AnnData.
mtb.cite
¶
Return citation text for the benchmark and, optionally, the methods you ran.
| PARAMETERS | DESCRIPTION |
|---|---|
*methods
|
Method ids, one per argument or as one list;
type
|
fmt
|
type
|
| RETURNS | DESCRIPTION |
|---|---|
str
|
The benchmark entry first, then one entry per method in the order given. |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
|
KeyError
|
Unknown method id; the message suggests a close match, if any. |
TypeError
|
Several ids are given and one is not a string. |
Examples
>>> import multibench as mtb
>>> print(mtb.cite()) # the benchmark only
>>> print(mtb.cite("Matilda", "MOFA2"))
>>> res = mtb.load_batch("out/")
>>> ran = res.summary.query("status not in ['SKIPPED', 'FAIL', 'TIMEOUT']")
>>> print(mtb.cite(ran.method, fmt="bibtex")) # methods that finished
Notes
Two spellings. Both forms return the same text, and cite(None) is
cite():
cite('Matilda', 'MOFA2')- one id per argument, like the CLImultibench cite Matilda MOFA2.cite(['Matilda', 'MOFA2'])- one list or tuple.
Formats.
"text"- oneAuthors. Title. Journal (year). https://doi.org/...line per entry."bibtex"- one@article{<id>_<year>, ...}per entry, separated by a blank line; the benchmark's key isscMultiBench_<year>.
Methods without a reference. A method without a verified reference
is emitted as a % <id>: no verified reference; see <repo_url>
comment (bibtex) or the same line without % (text).
See Also
mtb.method_info : carries the same reference, repo_url and version per method.
mtb.list_tasks
¶
List the task names that methods declare, sorted.
| RETURNS | DESCRIPTION |
|---|---|
list[str]
|
Sorted task names, the values |
Examples
>>> import multibench as mtb
>>> mtb.list_tasks() # ['batch', 'clustering', 'dimension_reduction']
>>> mtb.find_methods(task="dimension_reduction")
See Also
mtb.find_methods : filter methods by one of these tasks.
mtb.list_categories
¶
Return the valid category values with a plain-language description of each.
| RETURNS | DESCRIPTION |
|---|---|
dict
|
|
Examples
>>> import multibench as mtb
>>> mtb.list_categories()["vertical"]
'Several modalities measured in the same cells ...'
See Also
mtb.describe_layout : the file layout each category expects.