Prepare inputs¶
Find the files a method reads in a dataset folder. Both functions take
(dataset, category, method), in that order.
| Name | Summary |
|---|---|
mtb.inputs_for |
Return the input files a method reads from a dataset folder, by role. |
mtb.labels_for |
Return a dataset's cell-type label files, in the method's cell order. |
mtb.inputs_for
¶
inputs_for(
dataset: str,
category: str,
method: str,
*,
modalities: list[str] | set[str] | None = None,
data_path: Path | str | None = None,
check: bool | None = False,
) -> dict
Return the input files a method reads from a dataset folder, by role.
Paths are absolute, ready for mtb.run(inputs=...). Pass
check=True to verify the files before a long run.
| PARAMETERS | DESCRIPTION |
|---|---|
dataset
|
Dataset folder name under
type
|
category
|
Integration category:
type
|
method
|
Method id, e.g.
type
|
modalities
|
Modality tokens that pick the variant, e.g.
type
|
data_path
|
Data root that holds the dataset folders;
type
|
check
|
type
|
| RETURNS | DESCRIPTION |
|---|---|
dict
|
|
| RAISES | DESCRIPTION |
|---|---|
TypeError
|
|
KeyError
|
Unknown method, or no variant matches |
ValueError
|
Unknown category or modality token, or several variants fit
( |
ValueError
|
|
ValueError
|
|
FileNotFoundError
|
|
| WARNS | DESCRIPTION |
|---|---|
UserWarning
|
|
UserWarning
|
|
Examples
>>> import multibench as mtb
>>> mtb.inputs_for("D11", "vertical", "Matilda")
{'rna': '/abs/data/D11/rna.h5', 'adt': '/abs/data/D11/adt.h5',
'cty': '/abs/data/D11/cty.csv'}
>>> inp = mtb.inputs_for("D11", "vertical", "Matilda",
... modalities=["rna", "adt"], check=True)
>>> mtb.run("Matilda", "vertical", inputs=inp, out_dir="out/Matilda_D11")
>>> mtb.inputs_for("D28", "diagonal", "scBridge")
{'data_dir': '/abs/data/D28/'}
Notes
Variant selection. With modalities, the variant matching
(category, modalities) is used: exact role tokens first, then
atac standing for atac_gas / atac_peak when that leaves
exactly one variant. In the second case a representation token must
match what the method reads (method_info(m)['atac']): SCALEX with
["rna", "peak"] raises KeyError.
Without modalities, a category with one variant uses it. With several,
the dataset folder decides: the variant whose input files are all present
is used (Matilda on a rna.h5 + adt.h5 folder is its rna+adt variant).
When none or several qualify, mtb.AmbiguousVariantError (a
ValueError) lists the modality sets and the folder contents and asks
for modalities=.
Modality tokens. method_info(m)['supports'] lists each variant's
tokens.
proteinis another spelling ofadt;atacstands for either ATAC representation role;- the representation tokens
atac_peak(alsopeak,peaks) andatac_gas(alsogas,gene_activity) name what the method reads;method_info(m)['atac']says which one that is.
An unknown token raises ValueError naming the vocabulary.
Validation. A misspelt method id raises KeyError naming the
closest method id; an unknown category raises ValueError listing
the four.
File resolution. The dataset tree is flat
(<data_path>/<dataset>/<file>). Each role resolves to the file present
in the folder: the role token or a known alias. Label roles look for
.csv first. When nothing matches, the role falls back to <role>.h5
(<role>.csv for a label role), a path that does not exist (see
check).
Label files: the cty role of paired data reads cty.csv. An older
paired folder may name it rna_cty.csv; that file is read when the
folder has neither cty.csv nor atac_cty.csv.
ATAC files: vertical methods read atac.h5; method_info(m)['atac']
says whether it must hold peaks or gene activity. Diagonal methods read
atac_peak.h5 (peaks) and atac_gas.h5 (gene activity). Mosaic
methods read atac<i>.h5 (peaks). peak.h5, and atac.h5 for gene
activity, are accepted as older names, and so is atac_peak<i>.h5.
Numbered files: an unnumbered role (rna) also takes rna1.h5 when
the folder holds no rna2.h5. A folder with rna1.h5 and
rna2.h5 is per batch: vertical and diagonal roles stay unresolved,
and the error says to export without batch=.
A data_dir role (scBridge) resolves to the first of
<dataset>/processed/ and the dataset folder that holds a .h5ad
file; when neither does, to processed/ if that folder exists, else to
the dataset folder itself.
Absolute paths. Every returned path is absolute (a relative
data_path is resolved against the current directory), and a
data_dir value ends with the path separator. mtb.run executes the
method with cwd=out_dir, where a relative path would point at the
wrong place.
The check argument.
False(default): the best-effort paths, with no warning.None: the same paths, plus aUserWarninglisting the missing ones.True:FileNotFoundErrorfor a missing input, plus the content checks below.
The missing-file error or warning names an ATAC-family sibling that is
present, e.g. atac_peak.h5 when a vertical variant reads atac.h5,
and says when the folder holds per-batch files.
Content checks (check=True). The same checks mtb.scan
reports per row as files_ok / files_reason:
- orientation:
ValueErrorwhenmatrix/datais stored cells x features; - label length:
ValueErrorwhen a label CSV has a different number of rows than the modality file it labels, including the numberedcty<i>.csvof a cross/mosaic batch and the diagonalrna_cty.csv/atac_cty.csv(read by every evaluation, even when the method does not take them); data_dircontent:FileNotFoundErrorwhen a file scBridge names insidedata_dir(rna.h5,atac_gas.h5, the two label CSVs) is absent;- same cells:
ValueErrorwhen Seurat_v5'srna.h5andatac_peak.h5hold different cells. Seurat_v5 builds its paired bridge from these two files; - ATAC cell order (diagonal):
ValueErrorwhen theatac_gas.h5a method reads holds other cells thanatac_peak.h5, or lists them in another order.atac_cty.csvfollowsatac_peak.h5. Barcodes that differ only in a-1/-2suffix count as the same cell.
Dataset name case. A spelling that differs from the folder only in
case ('d52' for D52 on a case-insensitive filesystem) is replaced
by the on-disk spelling, with a UserWarning.
Argument order. (dataset, category, method), the same order
mtb.scan, mtb.run_all and mtb.labels_for use.
See Also
mtb.labels_for : the cell-type label CSVs of the same dataset, in the method's cell order.
mtb.run : consumes the returned dict as inputs=.
mtb.scan : the same resolution for every method at once, with reasons.
mtb.describe_layout : the folder layout these roles resolve against.
mtb.labels_for
¶
labels_for(
dataset: str,
category: str | None = None,
method: str | None = None,
*,
modalities: list[str] | set[str] | None = None,
data_path: Path | str | None = None,
check: bool | None = None,
) -> dict
Return a dataset's cell-type label files, in the method's cell order.
Hand the dict to mtb.evaluate(labels=...) as is. It matches an
embedding only when the embedding's cell order is the dict's order, so
give category and method for a method-specific order.
| PARAMETERS | DESCRIPTION |
|---|---|
dataset
|
Dataset folder name under
type
|
category
|
Integration category:
type
|
method
|
Method id; with
type
|
modalities
|
Modality tokens that pick one of several variants; used only with
type
|
data_path
|
Data root that holds the dataset folders;
type
|
check
|
Vertical or diagonal
type
|
| RETURNS | DESCRIPTION |
|---|---|
dict
|
|
| RAISES | DESCRIPTION |
|---|---|
TypeError
|
|
FileNotFoundError
|
|
ValueError
|
Unknown |
KeyError
|
Unknown |
| WARNS | DESCRIPTION |
|---|---|
UserWarning
|
|
UserWarning
|
|
Examples
>>> import multibench as mtb
>>> # diagonal: RNA cells first, then ATAC
>>> mtb.labels_for("D28")
{'rna_cty': '/abs/data/D28/rna_cty.csv',
'atac_cty': '/abs/data/D28/atac_cty.csv'}
>>> # StabMap: its reference batch first
>>> mtb.labels_for("D52", "cross", "StabMap")
{'cty3': '/abs/data/D52/cty3.csv', 'cty1': '/abs/data/D52/cty1.csv',
'cty2': '/abs/data/D52/cty2.csv'}
>>> inp = mtb.inputs_for("D52", "cross", "StabMap")
>>> res = mtb.run("StabMap", "cross", inputs=inp, out_dir="out/StabMap_D52")
>>> mtb.evaluate(res.output, labels=mtb.labels_for("D52", "cross", "StabMap"))
Notes
Which files. The benchmark stores cell-type labels as *cty*.csv in
the flat dataset folder, under dataset-specific names (cty.csv,
rna_cty.csv, cty1.csv ...). All of them are returned except the
tool-specific *_scjoint* reformats.
One exception: with category and method, a variant that reads
fewer numbered batches than the folder holds gets only the label files
of its batches. UINMF's cross variant reads batches 1 and 2, so on
D52 the dict holds cty1 and cty2.
An older paired folder may hold rna_cty.csv instead of cty.csv.
With category and method, a variant that reads cty gets that
file under the key cty, as mtb.inputs_for returns it.
Default order. Without category and method, the order is
not alphabetical:
cty(one file, cells already paired) first;- numbered
cty1, cty2, ..., cty10, in numeric order (batch order); - modality-named files in the canonical modality order rna, adt, atac
(
peak_ctycounts as atac) - the diagonal methods other than uniPort and Seurat_v5 emit the RNA cells first, then the ATAC cells; - any other
*cty*file, alphabetically, last.
Per-method order. With category and method, each file goes
where the cells it labels sit in that variant's output: the order of its
inputs, unless the method's cell order differs. uniPort and Seurat_v5
put their ATAC cells before their RNA cells.
StabMap uses a fixed reference batch: batch 3 in cross, batch 1 in
mosaic (method_info('StabMap')['supports'][i]['reference_batch']).
Its cell order starts with that batch (cty3, cty1, cty2 on D52).
Number the donor you want as reference accordingly.
The variant is chosen as mtb.inputs_for does: by modalities=, else
the category's only variant or the one whose files the folder holds. If
the choice stays ambiguous, or no label file pairs with the variant's
roles, the default order is used; labels_for does not raise
mtb.AmbiguousVariantError.
Passing it to evaluate. The dict is a dict subclass that remembers
its order, so mtb.evaluate(labels=...) takes it as is. A copy
(dict(d)) or a dict whose keys were reordered goes in as is only in
the default order; otherwise name the order with evaluate's
label_order=. list(labels_for(ds).values()) is the same files as a
list, in the same order, which evaluate also accepts.
Validation. category and method are validated whenever given:
ValueError listing the four categories on a typo, KeyError with a
did-you-mean hint for a method. Either one alone changes nothing.
Per-batch folders. A folder written with export_dataset(batch=...)
holds rna1.h5, rna2.h5 ... and cty1.csv, cty2.csv ....
Vertical and diagonal methods read one rna.h5 and one label file.
With such a category, check=None warns and check=True raises.
Export without batch= and pass the batch column to
mtb.run_all(batch=...) or mtb.evaluate(batch=...).
Paths and names. Paths are absolute. A dataset spelling that differs
from the folder only in case ('d52' for D52) is replaced by the
on-disk spelling, with a UserWarning. The positional order is
(dataset, category, method), like mtb.inputs_for, mtb.scan and
mtb.run_all.
See Also
mtb.inputs_for : the modality files of the same dataset.
mtb.evaluate : scores an embedding against these labels.