Skip to content

Evaluate a run

mtb.evaluate scores an embedding against cell-type labels and returns the scIB metrics. An embedding is a table with one row per cell, like adata.obsm["X_pca"].

Summary

import multibench as mtb

metrics = mtb.evaluate(
    result.output,                        # the embedding from mtb.run
    labels=mtb.labels_for("D11")["cty"],  # one label per cell, in row order
    metrics="clustering",
)
print(metrics)

Choose the metrics

metrics= takes a family ("clustering", "batch", "all") or a list of metric codes. The default computes every metric your inputs allow except kBET.

labels = mtb.labels_for("D11")["cty"]

mtb.evaluate(result.output, labels=labels, metrics=["ASW"])
# my_clusters: your own clusters, one per cell
mtb.evaluate(result.output, labels=labels, clustering=my_clusters,
             metrics=["ARI", "NMI"])
Details

ARI, NMI and iF1 need scIB's Leiden resolution sweep. The sweep clusters the embedding at ten resolutions from 0.2 to 2.0 and keeps the clustering with the highest NMI against the labels. A list without these three metrics skips it.

clustering= replaces the sweep for ARI and NMI. iF1 still runs the sweep unless you leave iF1 out.

kBET is computed only when a list names it, and can take hours on large datasets.


Inputs

Pass the embedding and one cell-type label per cell, in the embedding's row order. Both can be paths or in-memory objects. The reference lists the accepted forms.

# embedding and labels in one AnnData
mtb.evaluate(adata, labels="celltype", obsm="X_pca")
Details

With an AnnData, labels and batch can name .obs columns, and obsm= names the embedding. A MuData works the same way, with the joint embedding in its .obsm.

A Series indexed by cell id is aligned by id when the output has cell ids: an AnnData, an .h5ad file, or a DataFrame indexed by cell id. A CSV whose first column holds the cell ids is aligned the same way. When the output is a bare array, the Series and the CSV are matched by position, with a warning. Plain arrays, lists and other label files are matched by position.

Batch metrics

Batch metrics also need one batch per cell. Pass batch=, or the mtb.labels_for dict of a multi-batch dataset: each of its files is then a batch. Give labels_for the category and the method, so the files follow the rows of the method's output. A wrong order gives wrong scores without an error.

# a StabMap run on D52; labels_for lists cty3, cty1, cty2 for it
mtb.evaluate(
    result.output,
    labels=mtb.labels_for("D52", "cross", "StabMap"),
    metrics="all",
)
Details

Without a method, labels_for uses the default order: numbered files ascending, and rna before adt before atac. StabMap puts its reference batch first, which is cty3 on D52. What run returns lists the other methods with their own order.

evaluate takes a labels_for dict as it is. A dict you build yourself must follow the default order, or label_order= must list its keys in the order of the rows.

Without batch=, run_all uses each cell's label file as its batch, so batches inside one file, such as donors, are not scored. Pass batch= to run_all, or re-score saved outputs with the real batch vector. No method is re-run:

res = mtb.load_batch("out/")
rescored = res.rescore(batch=my_batch_vector)

The result

evaluate returns a DataFrame indexed by metric name, with one Value column. A metric that cannot run on your machine is NaN, with a warning. Save the scores with mtb.to_long.

mine = mtb.to_long(metrics, method="MyMethod", dataset="D11",
                   category="vertical")
mine.to_csv("mine.csv", index=False)
Details

A long table has one row per method, dataset and metric, like the tables mtb.load_results returns. Its source column reads user. Add your own method concatenates it with a stored table.

metrics.to_csv("metric.csv") writes the wide shape of the benchmark's metric.csv. That file loses the scoring record: to_long on it sets scored_with to unknown, with a warning.

What each metric measures

All metrics are scaled so that higher is better. Most lie between 0 and 1. ARI can be slightly below 0, which means a random clustering.

Metric Family What it measures
ARI clustering How well the Leiden clusters match the cell types
NMI clustering How much the Leiden clusters tell about the cell types
ASW clustering How well cells of one type separate from other types
iASW clustering The same, for each cell type against all other cells
iF1 clustering How well one Leiden cluster captures each cell type
cLISI clustering How much each cell's neighbours share its cell type
ASW_batch batch How well the batches mix within each cell type
GC batch How much of each cell type stays in one connected group of neighbours
iLISI batch How many batches appear among each cell's neighbours
kBET batch How often a neighbourhood holds the overall batch mix

mtb.catalog.metrics() describes each metric in one line. evaluate's Notes give the exact computation.

Details

cLISI and iLISI need scib's compiled LISI helper. Where the binary scib ships cannot run, evaluate rebuilds it once if a C++ compiler (g++, c++ or clang++) is on PATH. Otherwise the warning prints the build command.

kBET needs R through rpy2. Otherwise kBET is NaN.


Are my numbers comparable?

To compare with the stored tables, published or re-run, use the Leiden backend they were scored with and the dataset's own labels:

mtb.config.DEFAULT.leiden_flavor = "leidenalg"
metrics = mtb.evaluate(result.output, labels=mtb.labels_for("D11"))

The default igraph backend is faster but can move ARI by up to about 0.1. Most methods' scores vary slightly between runs. Compare ranks.

Details

mtb.config.DEFAULT.leiden_flavor takes "igraph", the default, or "leidenalg", the backend scib itself uses. Both stored sources were scored with leidenalg. The re-run tables were scored by evaluate from multibench 0.2.1.

iASW and iF1 score every cell type, whereas scib's default scores only types confined to few batches. The re-run tables use the same rule.

metrics.attrs records how the scores were computed: the Leiden backend, where the clusters came from, and the multibench and scib versions. mtb.to_long keeps the first three in its scored_with column, such as leidenalg/sweep/0.3.2.


Next: draw it next to the stored tables in Load and plot results.