Skip to content

Prepare inputs

Find the files a method reads in a dataset folder. Both functions take (dataset, category, method), in that order.

Name Summary
mtb.inputs_for Return the input files a method reads from a dataset folder, by role.
mtb.labels_for Return a dataset's cell-type label files, in the method's cell order.

mtb.inputs_for

inputs_for(
    dataset: str,
    category: str,
    method: str,
    *,
    modalities: list[str] | set[str] | None = None,
    data_path: Path | str | None = None,
    check: bool | None = False,
) -> dict

Return the input files a method reads from a dataset folder, by role.

Paths are absolute, ready for mtb.run(inputs=...). Pass check=True to verify the files before a long run.

PARAMETERS DESCRIPTION
dataset

Dataset folder name under data_path, e.g. "D11".

type str

category

Integration category: vertical, diagonal, mosaic or cross.

type str

method

Method id, e.g. "Matilda"; see mtb.list_methods().

type str

modalities

Modality tokens that pick the variant, e.g. ["rna", "adt"]; None = the files in the dataset folder decide.

type list[str] | set[str] | None default None

data_path

Data root that holds the dataset folders; None = mtb.config.DEFAULT.data_path.

type Path | str | None default None

check

False = no checks; True = raise on a missing or malformed input; None = warn about missing files.

type bool | None default False

RETURNS DESCRIPTION
dict

{role: absolute path}, one entry per input role of the selected variant (data_dir for the directory-input methods).

RAISES DESCRIPTION
TypeError

category and method were passed in swapped order.

KeyError

Unknown method, or no variant matches category and modalities.

ValueError

Unknown category or modality token, or several variants fit (mtb.AmbiguousVariantError).

ValueError

check=True: a transposed matrix, or label rows differ from the cell count.

ValueError

check=True: files that must hold the same cells, in one order, do not.

FileNotFoundError

check=True: an input file is missing, or data_dir lacks a file the method names.

WARNS DESCRIPTION
UserWarning

dataset matches a folder only up to letter case.

UserWarning

check=None and some resolved files do not exist.

Examples

>>> import multibench as mtb
>>> mtb.inputs_for("D11", "vertical", "Matilda")
{'rna': '/abs/data/D11/rna.h5', 'adt': '/abs/data/D11/adt.h5',
 'cty': '/abs/data/D11/cty.csv'}
>>> inp = mtb.inputs_for("D11", "vertical", "Matilda",
...                      modalities=["rna", "adt"], check=True)
>>> mtb.run("Matilda", "vertical", inputs=inp, out_dir="out/Matilda_D11")
>>> mtb.inputs_for("D28", "diagonal", "scBridge")
{'data_dir': '/abs/data/D28/'}
Notes

Variant selection. With modalities, the variant matching (category, modalities) is used: exact role tokens first, then atac standing for atac_gas / atac_peak when that leaves exactly one variant. In the second case a representation token must match what the method reads (method_info(m)['atac']): SCALEX with ["rna", "peak"] raises KeyError.

Without modalities, a category with one variant uses it. With several, the dataset folder decides: the variant whose input files are all present is used (Matilda on a rna.h5 + adt.h5 folder is its rna+adt variant). When none or several qualify, mtb.AmbiguousVariantError (a ValueError) lists the modality sets and the folder contents and asks for modalities=.

Modality tokens. method_info(m)['supports'] lists each variant's tokens.

  • protein is another spelling of adt;
  • atac stands for either ATAC representation role;
  • the representation tokens atac_peak (also peak, peaks) and atac_gas (also gas, gene_activity) name what the method reads; method_info(m)['atac'] says which one that is.

An unknown token raises ValueError naming the vocabulary.

Validation. A misspelt method id raises KeyError naming the closest method id; an unknown category raises ValueError listing the four.

File resolution. The dataset tree is flat (<data_path>/<dataset>/<file>). Each role resolves to the file present in the folder: the role token or a known alias. Label roles look for .csv first. When nothing matches, the role falls back to <role>.h5 (<role>.csv for a label role), a path that does not exist (see check).

Label files: the cty role of paired data reads cty.csv. An older paired folder may name it rna_cty.csv; that file is read when the folder has neither cty.csv nor atac_cty.csv.

ATAC files: vertical methods read atac.h5; method_info(m)['atac'] says whether it must hold peaks or gene activity. Diagonal methods read atac_peak.h5 (peaks) and atac_gas.h5 (gene activity). Mosaic methods read atac<i>.h5 (peaks). peak.h5, and atac.h5 for gene activity, are accepted as older names, and so is atac_peak<i>.h5.

Numbered files: an unnumbered role (rna) also takes rna1.h5 when the folder holds no rna2.h5. A folder with rna1.h5 and rna2.h5 is per batch: vertical and diagonal roles stay unresolved, and the error says to export without batch=.

A data_dir role (scBridge) resolves to the first of <dataset>/processed/ and the dataset folder that holds a .h5ad file; when neither does, to processed/ if that folder exists, else to the dataset folder itself.

Absolute paths. Every returned path is absolute (a relative data_path is resolved against the current directory), and a data_dir value ends with the path separator. mtb.run executes the method with cwd=out_dir, where a relative path would point at the wrong place.

The check argument.

  • False (default): the best-effort paths, with no warning.
  • None: the same paths, plus a UserWarning listing the missing ones.
  • True: FileNotFoundError for a missing input, plus the content checks below.

The missing-file error or warning names an ATAC-family sibling that is present, e.g. atac_peak.h5 when a vertical variant reads atac.h5, and says when the folder holds per-batch files.

Content checks (check=True). The same checks mtb.scan reports per row as files_ok / files_reason:

  • orientation: ValueError when matrix/data is stored cells x features;
  • label length: ValueError when a label CSV has a different number of rows than the modality file it labels, including the numbered cty<i>.csv of a cross/mosaic batch and the diagonal rna_cty.csv / atac_cty.csv (read by every evaluation, even when the method does not take them);
  • data_dir content: FileNotFoundError when a file scBridge names inside data_dir (rna.h5, atac_gas.h5, the two label CSVs) is absent;
  • same cells: ValueError when Seurat_v5's rna.h5 and atac_peak.h5 hold different cells. Seurat_v5 builds its paired bridge from these two files;
  • ATAC cell order (diagonal): ValueError when the atac_gas.h5 a method reads holds other cells than atac_peak.h5, or lists them in another order. atac_cty.csv follows atac_peak.h5. Barcodes that differ only in a -1 / -2 suffix count as the same cell.

Dataset name case. A spelling that differs from the folder only in case ('d52' for D52 on a case-insensitive filesystem) is replaced by the on-disk spelling, with a UserWarning.

Argument order. (dataset, category, method), the same order mtb.scan, mtb.run_all and mtb.labels_for use.

See Also

mtb.labels_for : the cell-type label CSVs of the same dataset, in the method's cell order.

mtb.run : consumes the returned dict as inputs=.

mtb.scan : the same resolution for every method at once, with reasons.

mtb.describe_layout : the folder layout these roles resolve against.

mtb.labels_for

labels_for(
    dataset: str,
    category: str | None = None,
    method: str | None = None,
    *,
    modalities: list[str] | set[str] | None = None,
    data_path: Path | str | None = None,
    check: bool | None = None,
) -> dict

Return a dataset's cell-type label files, in the method's cell order.

Hand the dict to mtb.evaluate(labels=...) as is. It matches an embedding only when the embedding's cell order is the dict's order, so give category and method for a method-specific order.

PARAMETERS DESCRIPTION
dataset

Dataset folder name under data_path, e.g. "D28".

type str

category

Integration category: vertical, diagonal, mosaic or cross; with method, selects the variant whose cell order is used. None = the default order.

type str | None default None

method

Method id; with category, orders the files in that variant's cell order. None = the default order.

type str | None default None

modalities

Modality tokens that pick one of several variants; used only with category and method.

type list[str] | set[str] | None default None

data_path

Data root that holds the dataset folders; None = mtb.config.DEFAULT.data_path.

type Path | str | None default None

check

Vertical or diagonal category on a per-batch folder: None warns, True raises, False = no check.

type bool | None default None

RETURNS DESCRIPTION
dict

{stem: absolute path} in the method's cell order, keyed by filename stem (cty, rna_cty, cty1 ...).

RAISES DESCRIPTION
TypeError

category and method swapped, or a path passed positionally as category.

FileNotFoundError

<data_path>/<dataset> does not exist.

ValueError

Unknown category; or check=True and a per-batch folder for vertical / diagonal.

KeyError

Unknown method, or it has no variant for category.

WARNS DESCRIPTION
UserWarning

dataset matches a folder only up to letter case.

UserWarning

check=None and a per-batch folder for a vertical or diagonal category.

Examples

>>> import multibench as mtb
>>> # diagonal: RNA cells first, then ATAC
>>> mtb.labels_for("D28")
{'rna_cty': '/abs/data/D28/rna_cty.csv',
 'atac_cty': '/abs/data/D28/atac_cty.csv'}
>>> # StabMap: its reference batch first
>>> mtb.labels_for("D52", "cross", "StabMap")
{'cty3': '/abs/data/D52/cty3.csv', 'cty1': '/abs/data/D52/cty1.csv',
 'cty2': '/abs/data/D52/cty2.csv'}
>>> inp = mtb.inputs_for("D52", "cross", "StabMap")
>>> res = mtb.run("StabMap", "cross", inputs=inp, out_dir="out/StabMap_D52")
>>> mtb.evaluate(res.output, labels=mtb.labels_for("D52", "cross", "StabMap"))
Notes

Which files. The benchmark stores cell-type labels as *cty*.csv in the flat dataset folder, under dataset-specific names (cty.csv, rna_cty.csv, cty1.csv ...). All of them are returned except the tool-specific *_scjoint* reformats.

One exception: with category and method, a variant that reads fewer numbered batches than the folder holds gets only the label files of its batches. UINMF's cross variant reads batches 1 and 2, so on D52 the dict holds cty1 and cty2.

An older paired folder may hold rna_cty.csv instead of cty.csv. With category and method, a variant that reads cty gets that file under the key cty, as mtb.inputs_for returns it.

Default order. Without category and method, the order is not alphabetical:

  1. cty (one file, cells already paired) first;
  2. numbered cty1, cty2, ..., cty10, in numeric order (batch order);
  3. modality-named files in the canonical modality order rna, adt, atac (peak_cty counts as atac) - the diagonal methods other than uniPort and Seurat_v5 emit the RNA cells first, then the ATAC cells;
  4. any other *cty* file, alphabetically, last.

Per-method order. With category and method, each file goes where the cells it labels sit in that variant's output: the order of its inputs, unless the method's cell order differs. uniPort and Seurat_v5 put their ATAC cells before their RNA cells.

StabMap uses a fixed reference batch: batch 3 in cross, batch 1 in mosaic (method_info('StabMap')['supports'][i]['reference_batch']). Its cell order starts with that batch (cty3, cty1, cty2 on D52). Number the donor you want as reference accordingly.

The variant is chosen as mtb.inputs_for does: by modalities=, else the category's only variant or the one whose files the folder holds. If the choice stays ambiguous, or no label file pairs with the variant's roles, the default order is used; labels_for does not raise mtb.AmbiguousVariantError.

Passing it to evaluate. The dict is a dict subclass that remembers its order, so mtb.evaluate(labels=...) takes it as is. A copy (dict(d)) or a dict whose keys were reordered goes in as is only in the default order; otherwise name the order with evaluate's label_order=. list(labels_for(ds).values()) is the same files as a list, in the same order, which evaluate also accepts.

Validation. category and method are validated whenever given: ValueError listing the four categories on a typo, KeyError with a did-you-mean hint for a method. Either one alone changes nothing.

Per-batch folders. A folder written with export_dataset(batch=...) holds rna1.h5, rna2.h5 ... and cty1.csv, cty2.csv .... Vertical and diagonal methods read one rna.h5 and one label file. With such a category, check=None warns and check=True raises. Export without batch= and pass the batch column to mtb.run_all(batch=...) or mtb.evaluate(batch=...).

Paths and names. Paths are absolute. A dataset spelling that differs from the folder only in case ('d52' for D52) is replaced by the on-disk spelling, with a UserWarning. The positional order is (dataset, category, method), like mtb.inputs_for, mtb.scan and mtb.run_all.

See Also

mtb.inputs_for : the modality files of the same dataset.

mtb.evaluate : scores an embedding against these labels.