Cross integration¶
Cross integration combines batches that all measure the same modalities. The task is to remove batch effects and keep the biological structure. Every cross method here reads RNA and ADT. For several 10x Multiome samples, use the vertical tutorial. This tutorial runs StabMap and sciPENN on D52, with 23,478 cells in three batches.
Methods run on Linux. On macOS or Windows, open this notebook in Colab.
1. Install¶
This cell installs multibench-sc. It keeps the numpy and pandas that are already installed, so Colab needs no restart.
Details
On Colab, choose a GPU runtime before you run the notebook: Runtime -> Change runtime type -> T4 GPU. On a CPU runtime, an environment that has a smaller CPU build gets that build, and training methods run slower.
import numpy, pandas
%pip install -q multibench-sc numpy=={numpy.__version__} pandas=={pandas.__version__}
Note: you may need to restart the kernel to use updated packages.
2. Download the data and the environments¶
mtb.data.fetch downloads D52 (179 MB) once. Each method runs in its own environment. mtb.env.install downloads the environments for StabMap and sciPENN, with no conda needed. The download is 1.4 GB on a computer without a GPU and 3.0 GB on one with a GPU.
Details
state is PACKED for an environment downloaded now and have for one that was already there. Without dry_run=False, mtb.env.install downloads nothing and returns the plan with its sizes.
Environments go to mtb.config.DEFAULT.envs_dir. To use another disk, set it before this cell.
From a terminal, multibench env install --methods StabMap,sciPENN --packed --run does the same. With --category cross instead of --methods, it installs every cross environment: 7 envs, 11.1 GB to download on a CPU host, 17.0 GB on a GPU host.
import pandas as pd
import multibench as mtb
mtb.data.fetch("D52")
downloading D52 (179 MB) from https://github.com/DSichang/scMultiBench/releases/download/data-v1/D52.tar.gz ...
PosixPath('/tmp/mbtut3/home/.cache/multibench/data')
METHODS = ["StabMap", "sciPENN"]
envs = mtb.env.install(METHODS, dry_run=False)
pd.DataFrame(envs)[["env", "methods", "state"]]
[env] downloading env_sciPENN (gpu build) 2.1 GB ...
[env] 0.2 GB of 2.1 GB (14 MB/s)
[env] 0.4 GB of 2.1 GB (15 MB/s)
[env] 0.6 GB of 2.1 GB (14 MB/s)
[env] 0.8 GB of 2.1 GB (15 MB/s)
[env] 1.0 GB of 2.1 GB (16 MB/s)
[env] 1.2 GB of 2.1 GB (15 MB/s)
[env] 1.5 GB of 2.1 GB (15 MB/s)
[env] 1.7 GB of 2.1 GB (15 MB/s)
[env] 1.9 GB of 2.1 GB (16 MB/s)
[env] 2.1 GB of 2.1 GB (16 MB/s)
[env] unpacking env_sciPENN -> /tmp/mbtut3/home/.cache/multibench/envs/env_sciPENN ...
[env] downloading scmb_r (single build) 0.9 GB ...
[env] 0.1 GB of 0.9 GB (14 MB/s)
[env] 0.2 GB of 0.9 GB (13 MB/s)
[env] 0.3 GB of 0.9 GB (14 MB/s)
[env] 0.4 GB of 0.9 GB (13 MB/s)
[env] 0.5 GB of 0.9 GB (14 MB/s)
[env] 0.6 GB of 0.9 GB (14 MB/s)
[env] 0.6 GB of 0.9 GB (14 MB/s)
[env] 0.7 GB of 0.9 GB (14 MB/s)
[env] 0.8 GB of 0.9 GB (15 MB/s)
[env] 0.9 GB of 0.9 GB (15 MB/s)
[env] unpacking scmb_r -> /tmp/mbtut3/home/.cache/multibench/envs/scmb_r ...
| env | methods | state | |
|---|---|---|---|
| 0 | env_sciPENN | [sciPENN] | PACKED |
| 1 | scmb_r | [StabMap] | PACKED |
3. Run the methods¶
run_all runs each method on D52 and scores its output with the scIB metrics. res.summary has one row per method with its status, run time and scores. Higher is better for every metric.
Details
Each method writes an embedding: a table of numbers with one row per cell. run_all scores it against the cell-type labels of D52.
The clustering metrics ARI, NMI, ASW, iASW, iF1 and cLISI measure how well the embedding separates the cell types. The batch metrics ASW_batch, GC and iLISI measure how well the batches mix. They appear only when the data has several batches.
res.failures says why a method failed. params={"Method": {"key": value}} sets a method's parameters, and mtb.params_for lists them. Many methods take none.
mtb.load_batch("out/D52") reloads these results later without running anything.
res = mtb.run_all("D52", "cross", methods=METHODS, out_dir="out/D52")
res.summary
[run_all] StabMap (cross/D52) ...
Fetching the method scripts from PYangLab/scMultiBench into /tmp/mbtut3/home/.cache/multibench/scMultiBench_ref ...
[run_all] -> CHAIN_OK (380.1s) ARI 0.683
[run_all] sciPENN (cross/D52) ...
[run_all] -> CHAIN_OK (190.6s) ARI 0.727
| method | status | run_sec | output_kind | emb_shape | n_tunable | label_order | label_order_confidence | batch_source | n_batches | ... | ASW | iASW | iF1 | cLISI | ASW_batch | GC | iLISI | label_order_note | caveat | reason | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | StabMap | CHAIN_OK | 380.1 | embedding | [23478, 50] | 0 | cty3.csv+cty1.csv+cty2.csv | 0.8252 | file_of_origin | 3 | ... | 0.5569 | 0.5656 | 0.6831 | 0.9939 | 0.9241 | 0.9326 | 0.5201 | None | ||
| 1 | sciPENN | CHAIN_OK | 190.6 | embedding | [23478, 512] | 1 | cty1.csv+cty2.csv+cty3.csv | 0.8416 | file_of_origin | 3 | ... | 0.6150 | 0.6296 | 0.7408 | 0.9978 | 0.8821 | 0.9369 | 0.3989 | None |
2 rows × 22 columns
4. Plot¶
res.plot() draws the scores as a bubble table. Circle size shows the rank within a column, and bigger is better. The fill compares a value with the other rows in the same column.
Details
The lightest fill is the lowest value in this figure, not zero.
Metrics are grouped by family: blue for dimension reduction and clustering, green for batch correction. Each family starts with an Overall bar. Its length and colour both show the family score.
A column whose rows all hold the same value is drawn grey, and the note under the figure names it.
res.plot()
5. Your own data¶
Your data needs raw RNA and ADT counts in one AnnData, with a cell type and a batch for each cell. mtb.io.export_dataset with batch= writes one set of files per batch, and run_all runs on that folder.
Details: export
For one AnnData per batch, call export_dataset once per batch with batch_index=N instead of batch=.
The folder name, MYCROSS, is the dataset name for run_all, and data_path is the folder that holds it.
mtb.scan("MYCROSS", "cross", data_path="mydata") checks the folder and the environments without running anything. Its reason column says what is missing.
overwrite=True replaces the files of an earlier run of this cell. Without it, export_dataset raises FileExistsError rather than replace a file.
Here one AnnData with 30% of D52's cells and a batch column takes the place of your data:
import anndata as ad
import scanpy as sc
d = mtb.config.DEFAULT.data_path / "D52"
batches = []
for b in (1, 2, 3):
a = mtb.io.read_canonical(d / f"rna{b}.h5")
a.obsm["protein"] = mtb.io.read_canonical(d / f"adt{b}.h5").to_df()
a.obs["celltype"] = pd.read_csv(d / f"cty{b}.csv")["x"].values
batches.append(a)
adata = ad.concat(batches, label="batch", keys=["1", "2", "3"], index_unique="-")
adata = sc.pp.subsample(adata, fraction=0.3, random_state=0, copy=True)
adata
AnnData object with n_obs × n_vars = 7043 × 10475
obs: 'celltype', 'batch'
obsm: 'protein'
Write the folder, then run the same methods on it:
mtb.io.export_dataset(adata, "mydata/MYCROSS", rna="X", adt="obsm:protein",
labels="obs:celltype", batch="obs:batch", category="cross",
overwrite=True)
mine = mtb.run_all("MYCROSS", "cross", methods=METHODS, data_path="mydata",
out_dir="out/MYCROSS")
mine.summary
[run_all] StabMap (cross/MYCROSS) ...
[run_all] -> CHAIN_OK (346.4s) ARI 0.708
[run_all] sciPENN (cross/MYCROSS) ...
[run_all] -> CHAIN_OK (150.2s) ARI 0.734
| method | status | run_sec | output_kind | emb_shape | n_tunable | label_order | label_order_confidence | batch_source | n_batches | ... | ASW | iASW | iF1 | cLISI | ASW_batch | GC | iLISI | label_order_note | caveat | reason | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | StabMap | CHAIN_OK | 346.4 | embedding | [7043, 50] | 0 | cty3.csv+cty1.csv+cty2.csv | 0.8359 | file_of_origin | 3 | ... | 0.5536 | 0.560 | 0.6519 | 0.9889 | 0.9113 | 0.9145 | 0.5886 | None | ||
| 1 | sciPENN | CHAIN_OK | 150.2 | embedding | [7043, 512] | 1 | cty1.csv+cty2.csv+cty3.csv | 0.8341 | file_of_origin | 3 | ... | 0.6111 | 0.611 | 0.7062 | 0.9952 | 0.8594 | 0.8948 | 0.5681 | None |
2 rows × 22 columns
mine.plot()
6. Stored scores¶
The package ships stored scores for 8 methods on D52. load_results reads them and mtb.plot.bubble draws them, without running anything.
Details
source="rerun" reads the package's own runs of the methods. source="published", the default, reads the published scIB tables, which hold 1 method for D52.
The stored scores used the leidenalg backend for Leiden clustering, and run_all uses igraph by default. The two backends can move ARI by up to about 0.1. To compare your runs with these scores, set mtb.config.DEFAULT.leiden_flavor = "leidenalg" before run_all.
long = mtb.load_results("cross", dataset="D52", source="rerun")
mtb.plot.bubble(long)
Troubleshooting¶
res.failures lists each method that failed or was skipped. Its error column ends with the method's error output.
Details
| symptom | fix |
|---|---|
env.install refuses on macOS or Windows |
methods run only on Linux: use Colab or a Linux machine |
| a method needs an NVIDIA GPU | choose a GPU runtime on Colab, or a machine with a GPU |
| a warning that values are not whole numbers | export raw counts, for example with rna="layer:counts" |
... matrix/data as cells x features |
the matrix is transposed: export it again with mtb.io.export_dataset |
| a method times out | raise timeout= in run_all |
low label_order_confidence |
several label files fit the cell count: check label_order_candidates in res.results |
| batch metrics use the wrong batches | res.rescore(batch=my_vector) scores again without running the methods |
Next steps¶
mtb.cite(METHODS)returns the citations for the benchmark and the methods you ran.- The other tutorials: vertical, diagonal, mosaic.
- The guides: run, evaluate, plot and discover methods.
- The interactive explorer has the full benchmark's rankings, with no install needed.