Data I/O (mtb.io)¶
Write your data as a dataset folder the methods can read, or convert single
matrices to and from the canonical .h5 format. Give raw counts: the methods
normalise the data themselves. A canonical .h5 holds one matrix as features
x cells, the layout the methods read.
| Name | Summary |
|---|---|
mtb.io.export_dataset |
Write an AnnData / MuData (or loose objects) as a canonical dataset folder. |
mtb.io.to_canonical |
Convert one matrix to a canonical scMultiBench .h5 and return its path. |
mtb.io.read_canonical |
Read a canonical .h5 back into an AnnData (cells x features). |
mtb.io.normalize_peak_names |
Copy a canonical .h5, rewriting ATAC peak names to chr:start-end. |
mtb.io.export_dataset
¶
export_dataset(
data,
dataset_dir: Path | str,
*,
rna="X",
adt=None,
atac=None,
atac_kind: str | None = None,
labels=None,
batch=None,
dtype: str = "float64",
compression: str | None = "gzip",
category: str | None = None,
adt_names: list | None = None,
batch_index: int | None = None,
overwrite: bool = False,
) -> Path
Write an AnnData / MuData (or loose objects) as a canonical dataset folder.
Produces the layout mtb.describe_layout documents (rna.h5,
adt.h5, ATAC files, cty.csv), so mtb.scan and
mtb.run_all work on your own data. Give raw counts.
| PARAMETERS | DESCRIPTION |
|---|---|
data
|
Object the selector strings refer to;
type
|
dataset_dir
|
Folder to create; its name is the dataset id for
type
|
rna
|
Raw-count matrix for RNA: a selector against
type
|
adt
|
Raw-count matrix for protein (ADT), same forms as
type
|
atac
|
ATAC matrix, same forms as
type
|
atac_kind
|
What
type
|
labels
|
Cell-type labels:
type
|
batch
|
Batch per cell, same forms as
type
|
dtype
|
Stored dtype of
type
|
compression
|
h5py compression filter, forwarded to
type
|
category
|
Integration category the folder is for (
type
|
adt_names
|
Protein names for the ADT matrix; they override any it carries and are needed when it has none.
type
|
batch_index
|
Write the whole object as batch
type
|
overwrite
|
type
|
| RETURNS | DESCRIPTION |
|---|---|
Path
|
|
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
Missing or conflicting arguments, or cells that do not pair across modalities (Notes). |
KeyError
|
A selector names a |
FileExistsError
|
A file this call would write exists and |
| WARNS | DESCRIPTION |
|---|---|
UserWarning
|
Values that are not whole numbers, feature-name problems, and the other cases in Notes. |
Examples
>>> import multibench as mtb
>>> mtb.io.export_dataset(adata, "data/MYCITE", rna="layer:counts",
... adt="obsm:protein", labels="obs:celltype")
>>> mtb.io.export_dataset(mdata, "data/MYMULTIOME", rna="rna", atac="atac",
... atac_kind="peak", labels="obs:celltype",
... category="vertical")
>>> mtb.io.export_dataset(rna, "data/LUNG", atac=atac, atac_kind="peak",
... labels="obs:cell_type", category="diagonal")
>>> mtb.scan("MYCITE", "vertical", data_path="data")
Notes
Modality arguments. rna, adt and atac each accept:
- a selector string against
data:'X','obsm:<key>','layer:<key>','mod:<name>','mod:<name>.obsm:<key>'or'mod:<name>.layer:<key>'; - an AnnData (its
.Xis written); - a DataFrame (index = cell barcodes, columns = features);
- a 2-D array or sparse matrix already in the master cell order.
An obsm selector takes its feature names from adt_names (ADT
only), else the columns of a DataFrame-valued entry, else
adata.uns['<key>_names'] (order in mtb.io.to_canonical).
All matrices are cells x features and are written transposed. With
data=None the default rna='X' means "no RNA": pass
rna=<AnnData>. Separate objects are paired by barcode:
mtb.io.export_dataset(rna_adata, "data/MYMULTI", atac=atac_adata,
atac_kind="peak", labels=rna_adata.obs["celltype"])
Feature filter. A selector may end with [<var column>=<value>].
Only the features whose adata.var column equals the value are
written. A 10x Multiome AnnData read with gex_only=False holds genes
and peaks in one X:
mtb.io.export_dataset(adata, "data/MYARC",
rna="X[feature_types=Gene Expression]",
atac="X[feature_types=Peaks]", atac_kind="peak",
labels="obs:celltype", category="vertical")
Raw counts. Methods normalise the data themselves; give raw counts. A
UserWarning says when sampled RNA, ADT or peak values are not whole
numbers (log-normalised data). rna='layer:counts' usually selects the
counts. Gene-activity scores are not checked.
MuData. A bare modality name selects mdata.mod[name].X
(rna='rna'); a full 'mod:<name>' selector works too. The default
rna='X' raises for a MuData: pass rna=<mod name> or rna=None.
labels / batch read 'obs:<col>' from the global mdata.obs.
'<mod>:<col>' reads mdata[mod].obs, then muon's copy
mdata.obs['<mod>:<col>']; 'mod:<mod>.obs:<col>' works too.
Master cell order. Every modality is re-indexed by barcode to one
order: data.obs_names, or the first modality's obs_names when
data is not given.
- the same barcodes in another order are reordered;
- a barcode set that differs raises
ValueErrornaming the strays; - non-unique barcodes cannot be matched by name, so the cells are paired
positionally with a
UserWarning(a different cell count still raises).
A bare array has no barcodes, so the order cannot be checked. Pass a DataFrame or AnnData when you are not sure of the order.
Labels. cty.csv is the single-column CSV the benchmark reads
(header x, one label per line); mtb.labels_for finds these files
again. A labels / batch Series is aligned by index to the master
barcodes; a plain RangeIndex of the right length is taken
positionally.
Diagonal. category='diagonal' pairs no barcodes: RNA and ATAC may
come from different cells. rna is written as rna.h5 and atac
as atac_peak.h5 or atac_gas.h5. labels is read from each
object's own .obs into rna_cty.csv and atac_cty.csv. A call
that writes one modality writes only that modality's label file.
For a MuData, labels='<col>' or '<mod>:<col>' reads <col>
from each modality's own .obs (mdata['rna'].obs for the RNA
cells, mdata['atac'].obs for the ATAC cells). When the two columns
have different names, export RNA and ATAC in two calls.
Batches. With batch, cells are split per batch value (sorted)
and numbered files are written: rna1.h5, rna2.h5 ...,
adt1.h5 ..., cty1.csv ... (the layout of the shipped D52). Only
cross and mosaic methods read numbered files. For a vertical folder,
export without batch and pass the batch column to
mtb.run_all(batch=...) or mtb.evaluate(batch=...).
batch_index=N writes the whole object as batch N. A mosaic
delivery of one file per batch (the D46 pattern: CITE-seq, Multiome,
RNA only) is one call per file:
kw = dict(labels="obs:cell_type", category="mosaic")
mtb.io.export_dataset(cite, "data/LAB", adt="obsm:protein",
batch_index=1, **kw)
mtb.io.export_dataset(multiome, "data/LAB", atac="obsm:atac",
atac_kind="peak", batch_index=2, **kw)
mtb.io.export_dataset(rna_only, "data/LAB", batch_index=3, **kw)
mtb.describe_layout('mosaic') lists the batch patterns the methods
read.
ATAC filenames. Vertical methods read atac.h5;
method_info(m)['atac'] says whether it must hold peaks or gene
activity. Diagonal methods read atac_peak.h5 (peaks) and
atac_gas.h5 (gene activity). Mosaic methods read atac<i>.h5
(peaks). peak.h5, and atac.h5 for gene activity, are accepted as
older names. By category:
'vertical':atac.h5for both kinds;'diagonal':atac_peak.h5oratac_gas.h5;'mosaic':atac<i>.h5;- no
category:atac_peak.h5plus a hard-linkedatac.h5(a copy when the file system refuses links), oratac_gas.h5only. Editing a hard-linked file edits both.
The representation is not recorded on disk. Check that
method_info(m)['atac'] is the kind exported. mtb.scan and
run_all skip a method whose file holds the other representation,
unless allow_atac_mismatch=True. mtb.run only warns, and the
method gives a wrong embedding.
Existing files. Every check runs before dataset_dir is created;
a failed call writes nothing. When a file the call would write already
exists, FileExistsError lists it; pass overwrite=True to replace
it. Other files in the folder are kept.
Errors. ValueError is raised for:
- nothing to export (
rna,adt,atacandlabelsallNone); atacwithout a validatac_kind, oratac_kindwithoutatac;- an unknown
category; - a selector string with
data=None, a malformed selector, or a feature filter that matches nothing; - a bare array whose cell order cannot be checked (no barcodes anywhere) or with the wrong row count;
- a modality whose barcodes differ from the master order (the message names the strays);
- a
labels/batchSeries missing cells, or a sequence of the wrong length; - a label that is missing (NaN, None or
''):mtb.evaluatewould score it as one more cell type; batchwithcategory='vertical'or'diagonal', or together withbatch_index;batch_indexthat is not a positive integer, or withoutcategory='mosaic'/'cross';category='diagonal'with labels but no matrix, or with a plain label sequence for both RNA and ATAC.
The KeyError message lists the names the object does have.
Warnings. UserWarning is emitted for:
- RNA, ADT or peak values that are not whole numbers;
- feature names that contradict
atac_kind, or none at all (feature_0..is written; for ADT passadt_names); - non-unique barcodes (paired positionally);
- a dense
matrix/dataover 1 GB; batchwithatacand nocategory: only mosaic methods read numbered ATAC files;category='mosaic'withatac_kind='gene_activity': every mosaic method reads peaks;category='mosaic'or'cross'withoutbatchorbatch_index.
See Also
mtb.io.to_canonical : one matrix -> one canonical file. mtb.describe_layout : the layout being produced and the role -> filename mapping. mtb.scan : confirm the folder is found and which methods can run on it.
mtb.io.to_canonical
¶
to_canonical(
src,
out: Path | str | None = None,
modality: str | None = None,
convert: bool = True,
*,
layer: str | None = None,
obsm: str | None = None,
mod: str | None = None,
dtype: str = "float64",
compression: str | None = "gzip",
block: int = 1024,
category: str | None = None,
feature_names: list | None = None,
) -> Path
Convert one matrix to a canonical scMultiBench .h5 and return its path.
The canonical layout is matrix/data (features x cells),
matrix/features and matrix/barcodes. A path that is already a
canonical .h5 is returned untouched, whatever the other arguments.
| PARAMETERS | DESCRIPTION |
|---|---|
src
|
Matrix to convert: an AnnData, a MuData (with
type
|
out
|
Output file, or a directory (existing or ending in
type
|
modality
|
type
|
convert
|
type
|
layer
|
Take the matrix from
type
|
obsm
|
Take the matrix from
type
|
mod
|
MuData modality to convert (
type
|
dtype
|
Stored dtype of
type
|
compression
|
h5py compression filter;
type
|
block
|
Number of features written per streaming step.
type
|
category
|
Integration category the file is for (
type
|
feature_names
|
Feature names that override
type
|
| RETURNS | DESCRIPTION |
|---|---|
Path
|
The canonical |
| RAISES | DESCRIPTION |
|---|---|
FileNotFoundError
|
|
ValueError
|
Invalid or conflicting arguments, or a non-canonical |
KeyError
|
|
ImportError
|
|
| WARNS | DESCRIPTION |
|---|---|
UserWarning
|
Values not whole numbers; feature names missing or contradicting |
UserWarning
|
|
Examples
>>> import multibench as mtb
>>> # writes data/MYCITE/rna.h5
>>> mtb.io.to_canonical(adata, "data/MYCITE/", modality="rna")
>>> mtb.io.to_canonical(adata, "data/MYCITE/", modality="adt", obsm="protein")
>>> mtb.io.to_canonical("counts.csv", "rna.h5") # explicit file name
>>> mtb.io.to_canonical(atac, "data/MYMULTI/", modality="peak",
... category="vertical")
Notes
Output location.
outa file path: written there;categorydoes not rename it;outa directory andmodalitygiven: the modality's canonical filename is appended (rna.h5,adt.h5, ...);out=Noneandmodalitygiven: that filename in the current directory;out=NonewithoutmodalityraisesValueError.
Modality aliases and checks. 'protein' -> 'adt',
'peak' -> 'atac_peak', 'gas' / 'gene_activity' ->
'atac_gas'. The ADT check is the obsm= rule under ADT
matrices. 'atac_peak' warns when fewer than half of the feature
names look like peaks (chr1:100-200, chr1_100_200,
chr1-100-200), 'atac_gas' when more than half do; plain
'atac' is not checked.
ADT matrices. With modality='adt' on an AnnData that has
obsm keys, obsm= or layer= must say where the protein matrix
is; otherwise adata.X (usually the RNA) would be written as
adt.h5, so it raises. layer and obsm are mutually exclusive.
The feature names of an obsm matrix come from the first of:
feature_names;- the columns, when the
obsmentry is a DataFrame; adata.uns[f'{obsm}_names'], when its length matches;feature_0.., with aUserWarning: passfeature_namesto keep the protein names.
ATAC filenames. Vertical methods read atac.h5;
method_info(m)['atac'] says whether it must hold peaks or gene
activity. Diagonal methods read atac_peak.h5 (peaks) and
atac_gas.h5 (gene activity). Mosaic methods read atac<i>.h5
(peaks). peak.h5, and atac.h5 for gene activity, are accepted as
older names.
With out a directory, category='vertical' writes atac.h5 for
either representation. Any other category, or none, names the file
after the representation: atac_peak.h5, atac_gas.h5, or
atac.h5 for the plain 'atac' role. For a mosaic batch, pass the
numbered file as out ('data/MYMOSAIC/atac2.h5').
Check the ATAC kind. Without category='vertical',
to_canonical(atac, d, modality='peak') writes atac_peak.h5, and
mtb.scan(d, 'vertical') finds no ATAC method runnable. The
representation is not recorded on disk: method_info(m)['atac'] says
which kind a method expects.
Gene activity next to peaks. With modality='gas' (no
category, or 'diagonal') and an atac_peak.h5 in the output
folder, the rows are written in that file's cell order, because
atac_cty.csv follows it. One line on stderr says so. Gene activity
for other cells gives a UserWarning.
A canonical .h5 is returned untouched: to reorder an existing
atac_gas.h5, pass mtb.io.read_canonical(path) as src.
Raw counts. Methods normalise the data themselves; give raw counts.
For modality 'rna', 'adt' or 'atac_peak', a
UserWarning says when sampled values are not whole numbers
(log-normalised data); layer='counts' usually holds the counts, and
for an ADT matrix taken with obsm=, another obsm key.
Streaming. Sparse matrices (CSR/CSC, in memory or inside an
.h5ad/.h5mu) are converted to CSC and written block features
at a time, without densifying the whole matrix. The defaults (gzip,
'float64') match the shipped benchmark files; any compression
enables chunking. A .csv / .tsv is read as cells x features, and
a non-numeric first column is used as the cell barcodes.
Size on disk. matrix/data is stored dense (features x cells x
itemsize): gzip shrinks the file, but every reader densifies it. A
UserWarning states the size when it exceeds 1 GB and suggests
filtering features or dtype='float32', which h5py / rhdf5 / hdf5r
read (as double in R).
Errors. ValueError is raised for:
out=Nonewithoutmodality;- an unknown
modalityorcategory; convert=Falseon a non-canonical input;- an
.h5withoutmatrix/data(the message lists the keys it does hold - a top-leveldatadataset is a method output, not an input); - an unsupported suffix;
modmissing for a MuData, or given for anything else;layerandobsmtogether;modality='adt'on an AnnData withobsmkeys but noobsm=/layer=;- a feature-name or barcode count that does not match the matrix.
The FileNotFoundError message names the path and the current
directory; the KeyError message lists the entries the object has.
See Also
mtb.io.export_dataset : a whole dataset folder (all modalities, labels, batches) in one call.
mtb.io.read_canonical : the inverse, canonical .h5 -> AnnData.
mtb.describe_layout : the folder layout and the role -> filename mapping.
mtb.io.read_canonical
¶
Read a canonical .h5 back into an AnnData (cells x features).
The inverse of mtb.io.to_canonical: matrix/data is transposed
back to cells x features, matrix/features become var_names and
matrix/barcodes become obs_names.
| PARAMETERS | DESCRIPTION |
|---|---|
path
|
Canonical
type
|
sparse
|
type
|
| RETURNS | DESCRIPTION |
|---|---|
AnnData
|
|
| RAISES | DESCRIPTION |
|---|---|
FileNotFoundError
|
|
Examples
>>> import multibench as mtb
>>> d = mtb.config.DEFAULT.data_path / "D11"
>>> rna = mtb.io.read_canonical(d / "rna.h5")
>>> rna.shape, rna.var_names[:3]
>>> adt = mtb.io.read_canonical(d / "adt.h5", sparse=False)
Notes
Where the demo data is. mtb.data.fetch("D11") downloads a demo
dataset into mtb.config.DEFAULT.data_path, not into the current
directory.
Memory. The whole matrix is read densely and transposed before the
sparsity decision, so memory peaks at the dense size (features x cells
x 8 bytes) even when the result is CSR. It is meant for the shipped
benchmark inputs and for what to_canonical writes, not for a
100k-cell peak matrix.
No validation. A file without matrix/data raises h5py's
KeyError. matrix/features and matrix/barcodes are optional;
without them AnnData's default names ('0', '1', ...) are kept.
See Also
mtb.io.to_canonical : write a canonical file from an AnnData / path.
mtb.io.normalize_peak_names
¶
Copy a canonical .h5, rewriting ATAC peak names to chr:start-end.
Signac's CreateChromatinAssay(sep=c(":","-")) (Seurat v3 and the
other Signac-based methods) expects that spelling. The source file is
never modified.
| PARAMETERS | DESCRIPTION |
|---|---|
src
|
Canonical
type
|
dst
|
Path of the copy; parent folders are created, an existing file is overwritten.
type
|
| RETURNS | DESCRIPTION |
|---|---|
Path
|
|
Examples
Notes
Recognised peaks. Peak ids come as chr_start_end,
chr-start-end or chr:start-end. A feature counts as a peak when
it reads <chr><sep><start><sep><end> with two integers, the first
separator one of _, -, : and the second _ or -; it
is rewritten as <chr>:<start>-<end>.
Everything else is kept. Other features (gene symbols, ids without
two numbers) pass through unchanged. matrix/data and
matrix/barcodes are copied as they are; only matrix/features is
replaced.
Inside mtb.run. mtb.run applies this itself for GLUE and
Seurat_v3, writing a per-run <role>_normpeaks.h5 copy next to the
converted inputs. Call it by hand only to prepare a file for a script
you run outside the wrapper.
See Also
mtb.io.to_canonical : write the canonical file in the first place.