Mosaic integration¶
Mosaic integration combines batches that share only some modalities. For example, a paired RNA + ATAC batch can link an RNA-only batch and an ATAC-only batch. Each method accepts one batch pattern, and every mosaic method reads ATAC as peaks. This tutorial runs StabMap and scMoMaT on D46, with 21,416 cells in three batches: RNA + ADT, RNA + ATAC, and RNA only.
Methods run on Linux. On macOS or Windows, open this notebook in Colab.
1. Install¶
This cell installs multibench-sc. It keeps the numpy and pandas that are already installed, so Colab needs no restart.
Details
On Colab, choose a GPU runtime before you run the notebook: Runtime -> Change runtime type -> T4 GPU. On a CPU runtime, an environment that has a smaller CPU build gets that build, and training methods run slower.
import numpy, pandas
%pip install -q multibench-sc numpy=={numpy.__version__} pandas=={pandas.__version__}
Note: you may need to restart the kernel to use updated packages.
2. Download the data and the environments¶
mtb.data.fetch downloads D46 (97 MB) once. Each method runs in its own environment. mtb.env.install downloads the environments for StabMap and scMoMaT, with no conda needed. The download is 1.8 GB on a computer without a GPU and 5.5 GB on one with a GPU.
Details
state is PACKED for an environment downloaded now and have for one that was already there. Without dry_run=False, mtb.env.install downloads nothing and returns the plan with its sizes.
Environments go to mtb.config.DEFAULT.envs_dir. To use another disk, set it before this cell.
From a terminal, multibench env install --methods StabMap,scMoMaT --packed --run does the same. With --category mosaic instead of --methods, it installs every mosaic environment: 6 envs, 15.6 GB to download on a CPU host, 20.8 GB on a GPU host.
import pandas as pd
import multibench as mtb
mtb.data.fetch("D46")
downloading D46 (97 MB) from https://github.com/DSichang/scMultiBench/releases/download/data-v1/D46.tar.gz ...
PosixPath('/tmp/mbtut3/home/.cache/multibench/data')
METHODS = ["StabMap", "scMoMaT"]
envs = mtb.env.install(METHODS, dry_run=False)
pd.DataFrame(envs)[["env", "methods", "state"]]
[env] downloading scmb_r (single build) 0.9 GB ...
[env] 0.1 GB of 0.9 GB (14 MB/s)
[env] 0.2 GB of 0.9 GB (18 MB/s)
[env] 0.3 GB of 0.9 GB (21 MB/s)
[env] 0.4 GB of 0.9 GB (23 MB/s)
[env] 0.5 GB of 0.9 GB (24 MB/s)
[env] 0.6 GB of 0.9 GB (23 MB/s)
[env] 0.6 GB of 0.9 GB (24 MB/s)
[env] 0.7 GB of 0.9 GB (25 MB/s)
[env] 0.8 GB of 0.9 GB (25 MB/s)
[env] 0.9 GB of 0.9 GB (24 MB/s)
[env] unpacking scmb_r -> /tmp/mbtut3/home/.cache/multibench/envs/scmb_r ...
[env] downloading scmb_torch (gpu build) 4.5 GB ...
[env] 0.5 GB of 4.5 GB (11 MB/s)
[env] 0.9 GB of 4.5 GB (11 MB/s)
[env] 1.4 GB of 4.5 GB (11 MB/s)
[env] 1.8 GB of 4.5 GB (10 MB/s)
[env] 2.3 GB of 4.5 GB (10 MB/s)
[env] 2.7 GB of 4.5 GB (10 MB/s)
[env] 3.2 GB of 4.5 GB (10 MB/s)
[env] 3.6 GB of 4.5 GB (10 MB/s)
[env] 4.1 GB of 4.5 GB (10 MB/s)
[env] 4.5 GB of 4.5 GB (10 MB/s)
[env] unpacking scmb_torch -> /tmp/mbtut3/home/.cache/multibench/envs/scmb_torch ...
| env | methods | state | |
|---|---|---|---|
| 0 | scmb_r | [StabMap] | PACKED |
| 1 | scmb_torch | [scMoMaT] | PACKED |
3. Run the methods¶
run_all runs each method on D46 and scores its output with the scIB metrics. res.summary has one row per method with its status, run time and scores. Higher is better for every metric.
Details
Each method writes an embedding: a table of numbers with one row per cell. run_all scores it against the cell-type labels of D46.
The clustering metrics ARI, NMI, ASW, iASW, iF1 and cLISI measure how well the embedding separates the cell types. The batch metrics ASW_batch, GC and iLISI measure how well the batches mix. They appear only when the data has several batches.
scMoMaT writes a graph instead of an embedding. run_all scores its UMAP, and its status reads CHAIN_OK_GRAPH_METHOD.
res.failures says why a method failed. params={"Method": {"key": value}} sets a method's parameters, and mtb.params_for lists them. Many methods take none.
mtb.load_batch("out/D46") reloads these results later without running anything.
res = mtb.run_all("D46", "mosaic", methods=METHODS, out_dir="out/D46")
res.summary
[run_all] StabMap (mosaic/D46) ...
Fetching the method scripts from PYangLab/scMultiBench into /tmp/mbtut3/home/.cache/multibench/scMultiBench_ref ...
[run_all] -> CHAIN_OK (1006.7s) ARI 0.585
[run_all] scMoMaT (mosaic/D46) ...
[run_all] -> CHAIN_OK_GRAPH_METHOD (1857.7s) ARI 0.282
| method | status | run_sec | output_kind | emb_shape | n_tunable | label_order | label_order_confidence | batch_source | n_batches | ... | ASW | iASW | iF1 | cLISI | ASW_batch | GC | iLISI | label_order_note | caveat | reason | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | StabMap | CHAIN_OK | 1006.7 | embedding | [21416, 50] | 0 | cty1.csv+cty2.csv+cty3.csv | 0.6700 | file_of_origin | 3 | ... | 0.5426 | 0.5313 | 0.4839 | 0.9945 | 0.8100 | 0.9557 | 0.0276 | None | ||
| 1 | scMoMaT | CHAIN_OK_GRAPH_METHOD | 1857.7 | graph | [21416, 2] | 0 | cty1.csv+cty2.csv+cty3.csv | 0.8411 | file_of_origin | 3 | ... | 0.4152 | 0.3471 | 0.3396 | 0.9882 | 0.4944 | 0.4239 | 0.1256 | None |
2 rows × 22 columns
4. Plot¶
res.plot() draws the scores as a bubble table. Circle size shows the rank within a column, and bigger is better. The fill compares a value with the other rows in the same column.
Details
The lightest fill is the lowest value in this figure, not zero.
Metrics are grouped by family: blue for dimension reduction and clustering, green for batch correction. Each family starts with an Overall bar. Its length and colour both show the family score.
A column whose rows all hold the same value is drawn grey, and the note under the figure names it.
res.plot()
5. Your own data¶
Your data needs raw counts, with one AnnData per batch and a cell type for each cell. mtb.io.export_dataset with batch_index= writes one batch per call. Number the batches to match a pattern that mtb.describe_layout("mosaic") lists, and give ATAC as peaks.
Details: export
mtb.describe_layout("mosaic") lists the batch patterns and the methods that accept each one. The command line writes one batch per call with multibench convert ... --category mosaic --batch-index N.
The folder name, MYMOSAIC, is the dataset name for run_all, and data_path is the folder that holds it.
mtb.scan("MYMOSAIC", "mosaic", data_path="mydata") checks the folder and the environments without running anything. Its reason column says what is missing.
overwrite=True replaces the files of an earlier run of this cell. Without it, export_dataset raises FileExistsError rather than replace a file.
Here the first 1,500 cells of each D46 batch take the place of your data:
d = mtb.config.DEFAULT.data_path / "D46"
n = 1500
rna = [mtb.io.read_canonical(d / f"rna{b}.h5")[:n] for b in (1, 2, 3)]
labels = [pd.read_csv(d / f"cty{b}.csv")["x"].values[:n] for b in (1, 2, 3)]
adt1 = mtb.io.read_canonical(d / "adt1.h5")[:n]
atac2 = mtb.io.read_canonical(d / "atac2.h5")[:n]
Write the folder, then run the same methods on it:
out = "mydata/MYMOSAIC"
# batch 1: RNA + ADT
mtb.io.export_dataset(rna[0], out, adt=adt1, labels=labels[0],
batch_index=1, category="mosaic", overwrite=True)
# batch 2: RNA + ATAC peaks
mtb.io.export_dataset(rna[1], out, atac=atac2, atac_kind="peak", labels=labels[1],
batch_index=2, category="mosaic", overwrite=True)
# batch 3: RNA only
mtb.io.export_dataset(rna[2], out, labels=labels[2],
batch_index=3, category="mosaic", overwrite=True)
mine = mtb.run_all("MYMOSAIC", "mosaic", methods=METHODS, data_path="mydata",
out_dir="out/MYMOSAIC")
mine.summary
[run_all] StabMap (mosaic/MYMOSAIC) ...
[run_all] -> CHAIN_OK (156.9s) ARI 0.560
[run_all] scMoMaT (mosaic/MYMOSAIC) ...
[run_all] -> CHAIN_OK_GRAPH_METHOD (661.8s) ARI 0.280
| method | status | run_sec | output_kind | emb_shape | n_tunable | label_order | label_order_confidence | batch_source | n_batches | ... | ASW | iASW | iF1 | cLISI | ASW_batch | GC | iLISI | label_order_note | caveat | reason | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | StabMap | CHAIN_OK | 156.9 | embedding | [4500, 50] | 0 | cty1.csv+cty2.csv+cty3.csv | 0.7351 | file_of_origin | 3 | ... | 0.5422 | 0.5200 | 0.4124 | 0.9896 | 0.7944 | 0.9434 | 0.1814 | None | ||
| 1 | scMoMaT | CHAIN_OK_GRAPH_METHOD | 661.8 | graph | [4500, 2] | 0 | cty1.csv+cty2.csv+cty3.csv | 0.8141 | file_of_origin | 3 | ... | 0.3500 | 0.3118 | 0.3466 | 0.9795 | 0.4974 | 0.5155 | 0.2775 | None |
2 rows × 22 columns
mine.plot()
6. Stored scores¶
The package ships stored scores for 4 methods on D45. load_results reads them and mtb.plot.bubble draws them, without running anything.
Details
There are no stored scores for D46. D45 is a larger mosaic dataset with another batch pattern.
source="rerun" reads the package's own runs of the methods. There is no published scIB table for mosaic, so source="published", the default, raises FileNotFoundError.
The stored scores used the leidenalg backend for Leiden clustering, and run_all uses igraph by default. The two backends can move ARI by up to about 0.1. To compare your runs with these scores, set mtb.config.DEFAULT.leiden_flavor = "leidenalg" before run_all.
long = mtb.load_results("mosaic", dataset="D45", source="rerun")
mtb.plot.bubble(long)
Troubleshooting¶
res.failures lists each method that failed or was skipped. Its error column ends with the method's error output.
Details
| symptom | fix |
|---|---|
env.install refuses on macOS or Windows |
methods run only on Linux: use Colab or a Linux machine |
| a method needs an NVIDIA GPU | choose a GPU runtime on Colab, or a machine with a GPU |
| a warning that values are not whole numbers | export raw counts, for example with rna="layer:counts" |
... matrix/data as cells x features |
the matrix is transposed: export it again with mtb.io.export_dataset |
| a method times out | raise timeout= in run_all |
low label_order_confidence |
several label files fit the cell count: check label_order_candidates in res.results |
| batch metrics use the wrong batches | res.rescore(batch=my_vector) scores again without running the methods |
Next steps¶
mtb.cite(METHODS)returns the citations for the benchmark and the methods you ran.- The other tutorials: vertical, diagonal, cross.
- The guides: run, evaluate, plot and discover methods.
- The interactive explorer has the full benchmark's rankings, with no install needed.