Skip to content

Data I/O (mtb.io)

Write your data as a dataset folder the methods can read, or convert single matrices to and from the canonical .h5 format. Give raw counts: the methods normalise the data themselves. A canonical .h5 holds one matrix as features x cells, the layout the methods read.

Name Summary
mtb.io.export_dataset Write an AnnData / MuData (or loose objects) as a canonical dataset folder.
mtb.io.to_canonical Convert one matrix to a canonical scMultiBench .h5 and return its path.
mtb.io.read_canonical Read a canonical .h5 back into an AnnData (cells x features).
mtb.io.normalize_peak_names Copy a canonical .h5, rewriting ATAC peak names to chr:start-end.

mtb.io.export_dataset

export_dataset(
    data,
    dataset_dir: Path | str,
    *,
    rna="X",
    adt=None,
    atac=None,
    atac_kind: str | None = None,
    labels=None,
    batch=None,
    dtype: str = "float64",
    compression: str | None = "gzip",
    category: str | None = None,
    adt_names: list | None = None,
    batch_index: int | None = None,
    overwrite: bool = False,
) -> Path

Write an AnnData / MuData (or loose objects) as a canonical dataset folder.

Produces the layout mtb.describe_layout documents (rna.h5, adt.h5, ATAC files, cty.csv), so mtb.scan and mtb.run_all work on your own data. Give raw counts.

PARAMETERS DESCRIPTION
data

Object the selector strings refer to; None when every modality is passed as an object. With a MuData, name each modality (rna='rna').

type (AnnData, MuData or None)

dataset_dir

Folder to create; its name is the dataset id for mtb.scan / mtb.run_all, and its parent their data_path.

type Path | str

rna

Raw-count matrix for RNA: a selector against data ('layer:counts', 'X[feature_types=Gene Expression]') or an object (Notes); None = no RNA.

type (str, AnnData, DataFrame, array or None) default 'X'

adt

Raw-count matrix for protein (ADT), same forms as rna (e.g. 'obsm:protein'); None = no ADT.

type (str, AnnData, DataFrame, array or None) default None

atac

ATAC matrix, same forms as rna (e.g. 'X[feature_types=Peaks]'); needs atac_kind. None = no ATAC.

type (str, AnnData, DataFrame, array or None) default None

atac_kind

What atac holds: 'peak' or 'gene_activity'; decides the ATAC filename.

type str or None default None

labels

Cell-type labels: 'obs:<col>' (a MuData's global .obs), '<mod>:<col>' (mdata[mod].obs), a Series indexed by barcode, or a sequence.

type (str, Series, sequence or None) default None

batch

Batch per cell, same forms as labels; splits the files per batch (rna1.h5, rna2.h5, ...) for mosaic or cross.

type (str, Series, sequence or None) default None

dtype

Stored dtype of matrix/data, forwarded to mtb.io.to_canonical.

type str default 'float64'

compression

h5py compression filter, forwarded to mtb.io.to_canonical.

type str or None default 'gzip'

category

Integration category the folder is for (vertical, diagonal, mosaic or cross); sets the ATAC and label filenames (Notes).

type str or None default None

adt_names

Protein names for the ADT matrix; they override any it carries and are needed when it has none.

type list or None default None

batch_index

Write the whole object as batch N (rna<N>.h5, cty<N>.csv); mosaic or cross, one call per batch file.

type int or None default None

overwrite

True = replace files already in dataset_dir; False = raise before writing anything.

type bool default False

RETURNS DESCRIPTION
Path

dataset_dir.

RAISES DESCRIPTION
ValueError

Missing or conflicting arguments, or cells that do not pair across modalities (Notes).

KeyError

A selector names a mod, obsm, layer, obs or var column that is absent.

FileExistsError

A file this call would write exists and overwrite is False.

WARNS DESCRIPTION
UserWarning

Values that are not whole numbers, feature-name problems, and the other cases in Notes.

Examples

>>> import multibench as mtb
>>> mtb.io.export_dataset(adata, "data/MYCITE", rna="layer:counts",
...                       adt="obsm:protein", labels="obs:celltype")
>>> mtb.io.export_dataset(mdata, "data/MYMULTIOME", rna="rna", atac="atac",
...                       atac_kind="peak", labels="obs:celltype",
...                       category="vertical")
>>> mtb.io.export_dataset(rna, "data/LUNG", atac=atac, atac_kind="peak",
...                       labels="obs:cell_type", category="diagonal")
>>> mtb.scan("MYCITE", "vertical", data_path="data")
Notes

Modality arguments. rna, adt and atac each accept:

  • a selector string against data: 'X', 'obsm:<key>', 'layer:<key>', 'mod:<name>', 'mod:<name>.obsm:<key>' or 'mod:<name>.layer:<key>';
  • an AnnData (its .X is written);
  • a DataFrame (index = cell barcodes, columns = features);
  • a 2-D array or sparse matrix already in the master cell order.

An obsm selector takes its feature names from adt_names (ADT only), else the columns of a DataFrame-valued entry, else adata.uns['<key>_names'] (order in mtb.io.to_canonical).

All matrices are cells x features and are written transposed. With data=None the default rna='X' means "no RNA": pass rna=<AnnData>. Separate objects are paired by barcode:

mtb.io.export_dataset(rna_adata, "data/MYMULTI", atac=atac_adata,
                      atac_kind="peak", labels=rna_adata.obs["celltype"])

Feature filter. A selector may end with [<var column>=<value>]. Only the features whose adata.var column equals the value are written. A 10x Multiome AnnData read with gex_only=False holds genes and peaks in one X:

mtb.io.export_dataset(adata, "data/MYARC",
                      rna="X[feature_types=Gene Expression]",
                      atac="X[feature_types=Peaks]", atac_kind="peak",
                      labels="obs:celltype", category="vertical")

Raw counts. Methods normalise the data themselves; give raw counts. A UserWarning says when sampled RNA, ADT or peak values are not whole numbers (log-normalised data). rna='layer:counts' usually selects the counts. Gene-activity scores are not checked.

MuData. A bare modality name selects mdata.mod[name].X (rna='rna'); a full 'mod:<name>' selector works too. The default rna='X' raises for a MuData: pass rna=<mod name> or rna=None.

labels / batch read 'obs:<col>' from the global mdata.obs. '<mod>:<col>' reads mdata[mod].obs, then muon's copy mdata.obs['<mod>:<col>']; 'mod:<mod>.obs:<col>' works too.

Master cell order. Every modality is re-indexed by barcode to one order: data.obs_names, or the first modality's obs_names when data is not given.

  • the same barcodes in another order are reordered;
  • a barcode set that differs raises ValueError naming the strays;
  • non-unique barcodes cannot be matched by name, so the cells are paired positionally with a UserWarning (a different cell count still raises).

A bare array has no barcodes, so the order cannot be checked. Pass a DataFrame or AnnData when you are not sure of the order.

Labels. cty.csv is the single-column CSV the benchmark reads (header x, one label per line); mtb.labels_for finds these files again. A labels / batch Series is aligned by index to the master barcodes; a plain RangeIndex of the right length is taken positionally.

Diagonal. category='diagonal' pairs no barcodes: RNA and ATAC may come from different cells. rna is written as rna.h5 and atac as atac_peak.h5 or atac_gas.h5. labels is read from each object's own .obs into rna_cty.csv and atac_cty.csv. A call that writes one modality writes only that modality's label file.

For a MuData, labels='<col>' or '<mod>:<col>' reads <col> from each modality's own .obs (mdata['rna'].obs for the RNA cells, mdata['atac'].obs for the ATAC cells). When the two columns have different names, export RNA and ATAC in two calls.

Batches. With batch, cells are split per batch value (sorted) and numbered files are written: rna1.h5, rna2.h5 ..., adt1.h5 ..., cty1.csv ... (the layout of the shipped D52). Only cross and mosaic methods read numbered files. For a vertical folder, export without batch and pass the batch column to mtb.run_all(batch=...) or mtb.evaluate(batch=...).

batch_index=N writes the whole object as batch N. A mosaic delivery of one file per batch (the D46 pattern: CITE-seq, Multiome, RNA only) is one call per file:

kw = dict(labels="obs:cell_type", category="mosaic")
mtb.io.export_dataset(cite, "data/LAB", adt="obsm:protein",
                      batch_index=1, **kw)
mtb.io.export_dataset(multiome, "data/LAB", atac="obsm:atac",
                      atac_kind="peak", batch_index=2, **kw)
mtb.io.export_dataset(rna_only, "data/LAB", batch_index=3, **kw)

mtb.describe_layout('mosaic') lists the batch patterns the methods read.

ATAC filenames. Vertical methods read atac.h5; method_info(m)['atac'] says whether it must hold peaks or gene activity. Diagonal methods read atac_peak.h5 (peaks) and atac_gas.h5 (gene activity). Mosaic methods read atac<i>.h5 (peaks). peak.h5, and atac.h5 for gene activity, are accepted as older names. By category:

  • 'vertical': atac.h5 for both kinds;
  • 'diagonal': atac_peak.h5 or atac_gas.h5;
  • 'mosaic': atac<i>.h5;
  • no category: atac_peak.h5 plus a hard-linked atac.h5 (a copy when the file system refuses links), or atac_gas.h5 only. Editing a hard-linked file edits both.

The representation is not recorded on disk. Check that method_info(m)['atac'] is the kind exported. mtb.scan and run_all skip a method whose file holds the other representation, unless allow_atac_mismatch=True. mtb.run only warns, and the method gives a wrong embedding.

Existing files. Every check runs before dataset_dir is created; a failed call writes nothing. When a file the call would write already exists, FileExistsError lists it; pass overwrite=True to replace it. Other files in the folder are kept.

Errors. ValueError is raised for:

  • nothing to export (rna, adt, atac and labels all None);
  • atac without a valid atac_kind, or atac_kind without atac;
  • an unknown category;
  • a selector string with data=None, a malformed selector, or a feature filter that matches nothing;
  • a bare array whose cell order cannot be checked (no barcodes anywhere) or with the wrong row count;
  • a modality whose barcodes differ from the master order (the message names the strays);
  • a labels / batch Series missing cells, or a sequence of the wrong length;
  • a label that is missing (NaN, None or ''): mtb.evaluate would score it as one more cell type;
  • batch with category='vertical' or 'diagonal', or together with batch_index;
  • batch_index that is not a positive integer, or without category='mosaic' / 'cross';
  • category='diagonal' with labels but no matrix, or with a plain label sequence for both RNA and ATAC.

The KeyError message lists the names the object does have.

Warnings. UserWarning is emitted for:

  • RNA, ADT or peak values that are not whole numbers;
  • feature names that contradict atac_kind, or none at all (feature_0.. is written; for ADT pass adt_names);
  • non-unique barcodes (paired positionally);
  • a dense matrix/data over 1 GB;
  • batch with atac and no category: only mosaic methods read numbered ATAC files;
  • category='mosaic' with atac_kind='gene_activity': every mosaic method reads peaks;
  • category='mosaic' or 'cross' without batch or batch_index.
See Also

mtb.io.to_canonical : one matrix -> one canonical file. mtb.describe_layout : the layout being produced and the role -> filename mapping. mtb.scan : confirm the folder is found and which methods can run on it.

mtb.io.to_canonical

to_canonical(
    src,
    out: Path | str | None = None,
    modality: str | None = None,
    convert: bool = True,
    *,
    layer: str | None = None,
    obsm: str | None = None,
    mod: str | None = None,
    dtype: str = "float64",
    compression: str | None = "gzip",
    block: int = 1024,
    category: str | None = None,
    feature_names: list | None = None,
) -> Path

Convert one matrix to a canonical scMultiBench .h5 and return its path.

The canonical layout is matrix/data (features x cells), matrix/features and matrix/barcodes. A path that is already a canonical .h5 is returned untouched, whatever the other arguments.

PARAMETERS DESCRIPTION
src

Matrix to convert: an AnnData, a MuData (with mod), or a path to .h5ad / .h5mu / .csv / .tsv (cells x features) / .loom / canonical .h5.

type (AnnData, MuData, Path or str)

out

Output file, or a directory (existing or ending in /) that gets the modality filename; None = that filename in the current directory (needs modality).

type Path | str | None default None

modality

'rna', 'adt', 'atac', 'atac_peak' or 'atac_gas'; names the file when out is a directory or None, and enables the ADT / ATAC checks (Notes).

type str | None default None

convert

False = never write: pass a canonical .h5 through, raise for anything else.

type bool default True

layer

Take the matrix from adata.layers[layer] instead of adata.X.

type str | None default None

obsm

Take the matrix from adata.obsm[obsm] (e.g. 'protein' for CITE-seq).

type str | None default None

mod

MuData modality to convert (mdata.mod[mod]); required for a MuData.

type str | None default None

dtype

Stored dtype of matrix/data; 'float32' halves the file (see Notes).

type str default 'float64'

compression

h5py compression filter; None = uncompressed.

type str | None default 'gzip'

block

Number of features written per streaming step.

type int default 1024

category

Integration category the file is for (vertical, diagonal, mosaic or cross); only changes ATAC filenames (Notes).

type str | None default None

feature_names

Feature names that override var_names / uns / DataFrame columns; the way to name a bare obsm array.

type list | None default None

RETURNS DESCRIPTION
Path

The canonical .h5 written, or src itself on passthrough.

RAISES DESCRIPTION
FileNotFoundError

src is a path that does not exist.

ValueError

Invalid or conflicting arguments, or a non-canonical .h5 (full list in Notes).

KeyError

layer, obsm or mod names an entry the object does not have.

ImportError

.h5mu without mudata, or .loom without loompy.

WARNS DESCRIPTION
UserWarning

Values not whole numbers; feature names missing or contradicting modality; size over 1 GB.

UserWarning

modality='gas' for other cells than the atac_peak.h5 next to out.

Examples

>>> import multibench as mtb
>>> # writes data/MYCITE/rna.h5
>>> mtb.io.to_canonical(adata, "data/MYCITE/", modality="rna")
>>> mtb.io.to_canonical(adata, "data/MYCITE/", modality="adt", obsm="protein")
>>> mtb.io.to_canonical("counts.csv", "rna.h5")  # explicit file name
>>> mtb.io.to_canonical(atac, "data/MYMULTI/", modality="peak",
...                     category="vertical")
Notes

Output location.

  • out a file path: written there; category does not rename it;
  • out a directory and modality given: the modality's canonical filename is appended (rna.h5, adt.h5, ...);
  • out=None and modality given: that filename in the current directory; out=None without modality raises ValueError.

Modality aliases and checks. 'protein' -> 'adt', 'peak' -> 'atac_peak', 'gas' / 'gene_activity' -> 'atac_gas'. The ADT check is the obsm= rule under ADT matrices. 'atac_peak' warns when fewer than half of the feature names look like peaks (chr1:100-200, chr1_100_200, chr1-100-200), 'atac_gas' when more than half do; plain 'atac' is not checked.

ADT matrices. With modality='adt' on an AnnData that has obsm keys, obsm= or layer= must say where the protein matrix is; otherwise adata.X (usually the RNA) would be written as adt.h5, so it raises. layer and obsm are mutually exclusive. The feature names of an obsm matrix come from the first of:

  1. feature_names;
  2. the columns, when the obsm entry is a DataFrame;
  3. adata.uns[f'{obsm}_names'], when its length matches;
  4. feature_0.., with a UserWarning: pass feature_names to keep the protein names.

ATAC filenames. Vertical methods read atac.h5; method_info(m)['atac'] says whether it must hold peaks or gene activity. Diagonal methods read atac_peak.h5 (peaks) and atac_gas.h5 (gene activity). Mosaic methods read atac<i>.h5 (peaks). peak.h5, and atac.h5 for gene activity, are accepted as older names.

With out a directory, category='vertical' writes atac.h5 for either representation. Any other category, or none, names the file after the representation: atac_peak.h5, atac_gas.h5, or atac.h5 for the plain 'atac' role. For a mosaic batch, pass the numbered file as out ('data/MYMOSAIC/atac2.h5').

Check the ATAC kind. Without category='vertical', to_canonical(atac, d, modality='peak') writes atac_peak.h5, and mtb.scan(d, 'vertical') finds no ATAC method runnable. The representation is not recorded on disk: method_info(m)['atac'] says which kind a method expects.

Gene activity next to peaks. With modality='gas' (no category, or 'diagonal') and an atac_peak.h5 in the output folder, the rows are written in that file's cell order, because atac_cty.csv follows it. One line on stderr says so. Gene activity for other cells gives a UserWarning. A canonical .h5 is returned untouched: to reorder an existing atac_gas.h5, pass mtb.io.read_canonical(path) as src.

Raw counts. Methods normalise the data themselves; give raw counts. For modality 'rna', 'adt' or 'atac_peak', a UserWarning says when sampled values are not whole numbers (log-normalised data); layer='counts' usually holds the counts, and for an ADT matrix taken with obsm=, another obsm key.

Streaming. Sparse matrices (CSR/CSC, in memory or inside an .h5ad/.h5mu) are converted to CSC and written block features at a time, without densifying the whole matrix. The defaults (gzip, 'float64') match the shipped benchmark files; any compression enables chunking. A .csv / .tsv is read as cells x features, and a non-numeric first column is used as the cell barcodes.

Size on disk. matrix/data is stored dense (features x cells x itemsize): gzip shrinks the file, but every reader densifies it. A UserWarning states the size when it exceeds 1 GB and suggests filtering features or dtype='float32', which h5py / rhdf5 / hdf5r read (as double in R).

Errors. ValueError is raised for:

  • out=None without modality;
  • an unknown modality or category;
  • convert=False on a non-canonical input;
  • an .h5 without matrix/data (the message lists the keys it does hold - a top-level data dataset is a method output, not an input);
  • an unsupported suffix;
  • mod missing for a MuData, or given for anything else;
  • layer and obsm together;
  • modality='adt' on an AnnData with obsm keys but no obsm= / layer=;
  • a feature-name or barcode count that does not match the matrix.

The FileNotFoundError message names the path and the current directory; the KeyError message lists the entries the object has.

See Also

mtb.io.export_dataset : a whole dataset folder (all modalities, labels, batches) in one call. mtb.io.read_canonical : the inverse, canonical .h5 -> AnnData. mtb.describe_layout : the folder layout and the role -> filename mapping.

mtb.io.read_canonical

read_canonical(path: Path | str, sparse: bool | None = None)

Read a canonical .h5 back into an AnnData (cells x features).

The inverse of mtb.io.to_canonical: matrix/data is transposed back to cells x features, matrix/features become var_names and matrix/barcodes become obs_names.

PARAMETERS DESCRIPTION
path

Canonical .h5 file.

type Path | str

sparse

True = CSR .X; False = dense ndarray; None = CSR when fewer than half of the entries are non-zero.

type bool | None default None

RETURNS DESCRIPTION
AnnData

.X as float, cells x features; names taken from the file when it has them.

RAISES DESCRIPTION
FileNotFoundError

path does not exist; the message names mtb.config.DEFAULT.data_path.

Examples

>>> import multibench as mtb
>>> d = mtb.config.DEFAULT.data_path / "D11"
>>> rna = mtb.io.read_canonical(d / "rna.h5")
>>> rna.shape, rna.var_names[:3]
>>> adt = mtb.io.read_canonical(d / "adt.h5", sparse=False)
Notes

Where the demo data is. mtb.data.fetch("D11") downloads a demo dataset into mtb.config.DEFAULT.data_path, not into the current directory.

Memory. The whole matrix is read densely and transposed before the sparsity decision, so memory peaks at the dense size (features x cells x 8 bytes) even when the result is CSR. It is meant for the shipped benchmark inputs and for what to_canonical writes, not for a 100k-cell peak matrix.

No validation. A file without matrix/data raises h5py's KeyError. matrix/features and matrix/barcodes are optional; without them AnnData's default names ('0', '1', ...) are kept.

See Also

mtb.io.to_canonical : write a canonical file from an AnnData / path.

mtb.io.normalize_peak_names

normalize_peak_names(src, dst)

Copy a canonical .h5, rewriting ATAC peak names to chr:start-end.

Signac's CreateChromatinAssay(sep=c(":","-")) (Seurat v3 and the other Signac-based methods) expects that spelling. The source file is never modified.

PARAMETERS DESCRIPTION
src

Canonical .h5 whose matrix/features holds the peak ids.

type Path | str

dst

Path of the copy; parent folders are created, an existing file is overwritten.

type Path | str

RETURNS DESCRIPTION
Path

dst.

Examples

>>> import multibench as mtb
>>> mtb.io.normalize_peak_names("data/MYMULTI/atac.h5", "tmp/atac.h5")
Notes

Recognised peaks. Peak ids come as chr_start_end, chr-start-end or chr:start-end. A feature counts as a peak when it reads <chr><sep><start><sep><end> with two integers, the first separator one of _, -, : and the second _ or -; it is rewritten as <chr>:<start>-<end>.

Everything else is kept. Other features (gene symbols, ids without two numbers) pass through unchanged. matrix/data and matrix/barcodes are copied as they are; only matrix/features is replaced.

Inside mtb.run. mtb.run applies this itself for GLUE and Seurat_v3, writing a per-run <role>_normpeaks.h5 copy next to the converted inputs. Call it by hand only to prepare a file for a script you run outside the wrapper.

See Also

mtb.io.to_canonical : write the canonical file in the first place.