Evaluate a run¶
mtb.evaluate scores an embedding against cell-type labels and returns the
scIB metrics. An embedding is a table with one row per cell, like
adata.obsm["X_pca"].
Summary
Choose the metrics¶
metrics= takes a family ("clustering",
"batch", "all") or a list of metric codes. The default computes every
metric your inputs allow except kBET.
labels = mtb.labels_for("D11")["cty"]
mtb.evaluate(result.output, labels=labels, metrics=["ASW"])
# my_clusters: your own clusters, one per cell
mtb.evaluate(result.output, labels=labels, clustering=my_clusters,
metrics=["ARI", "NMI"])
Details
ARI, NMI and iF1 need scIB's Leiden resolution sweep. The sweep clusters the embedding at ten resolutions from 0.2 to 2.0 and keeps the clustering with the highest NMI against the labels. A list without these three metrics skips it.
clustering= replaces the sweep for ARI and NMI. iF1 still runs the
sweep unless you leave iF1 out.
kBET is computed only when a list names it, and can take hours on large datasets.
Inputs¶
Pass the embedding and one cell-type label per cell, in the embedding's row order. Both can be paths or in-memory objects. The reference lists the accepted forms.
Details
With an AnnData, labels and batch can name .obs columns, and
obsm= names the embedding. A MuData works the same way, with the
joint embedding in its .obsm.
A Series indexed by cell id is aligned by id when the output has
cell ids: an AnnData, an .h5ad file, or a DataFrame indexed by cell
id. A CSV whose first column holds the cell ids is aligned the same
way. When the output is a bare array, the Series and the CSV are
matched by position, with a warning. Plain arrays, lists and other
label files are matched by position.
Batch metrics¶
Batch metrics also need one batch per cell. Pass batch=, or the
mtb.labels_for dict of a multi-batch dataset: each of its files is then
a batch. Give labels_for the category and the method, so the files
follow the rows of the method's output. A wrong order gives wrong scores
without an error.
# a StabMap run on D52; labels_for lists cty3, cty1, cty2 for it
mtb.evaluate(
result.output,
labels=mtb.labels_for("D52", "cross", "StabMap"),
metrics="all",
)
Details
Without a method, labels_for uses the default order: numbered files
ascending, and rna before adt before atac. StabMap puts its
reference batch first, which is cty3 on D52.
What run returns lists the other methods
with their own order.
evaluate takes a labels_for dict as it is. A dict you build
yourself must follow the default order, or label_order= must list
its keys in the order of the rows.
Without batch=, run_all uses each cell's label file as its batch,
so batches inside one file, such as donors, are not scored. Pass
batch= to run_all, or re-score saved outputs with the real batch
vector. No method is re-run:
The result¶
evaluate returns a DataFrame indexed by metric name, with one Value
column. A metric that cannot run on your machine is NaN, with a warning.
Save the scores with mtb.to_long.
mine = mtb.to_long(metrics, method="MyMethod", dataset="D11",
category="vertical")
mine.to_csv("mine.csv", index=False)
Details
A long table has one row per method, dataset and metric, like the
tables mtb.load_results returns. Its source column reads user.
Add your own method concatenates it
with a stored table.
metrics.to_csv("metric.csv") writes the wide shape of the
benchmark's metric.csv. That file loses the scoring record: to_long
on it sets scored_with to unknown, with a warning.
What each metric measures¶
All metrics are scaled so that higher is better. Most lie between 0 and 1. ARI can be slightly below 0, which means a random clustering.
| Metric | Family | What it measures |
|---|---|---|
| ARI | clustering | How well the Leiden clusters match the cell types |
| NMI | clustering | How much the Leiden clusters tell about the cell types |
| ASW | clustering | How well cells of one type separate from other types |
| iASW | clustering | The same, for each cell type against all other cells |
| iF1 | clustering | How well one Leiden cluster captures each cell type |
| cLISI | clustering | How much each cell's neighbours share its cell type |
| ASW_batch | batch | How well the batches mix within each cell type |
| GC | batch | How much of each cell type stays in one connected group of neighbours |
| iLISI | batch | How many batches appear among each cell's neighbours |
| kBET | batch | How often a neighbourhood holds the overall batch mix |
mtb.catalog.metrics()
describes each metric in one line.
evaluate's Notes
give the exact computation.
Details
cLISI and iLISI need scib's compiled LISI helper. Where the binary
scib ships cannot run, evaluate rebuilds it once if a C++ compiler
(g++, c++ or clang++) is on PATH. Otherwise the warning prints
the build command.
kBET needs R through rpy2. Otherwise kBET is NaN.
Are my numbers comparable?¶
To compare with the stored tables, published or re-run, use the Leiden backend they were scored with and the dataset's own labels:
mtb.config.DEFAULT.leiden_flavor = "leidenalg"
metrics = mtb.evaluate(result.output, labels=mtb.labels_for("D11"))
The default igraph backend is faster but can move ARI by up to about
0.1. Most methods' scores vary slightly between runs. Compare ranks.
Details
mtb.config.DEFAULT.leiden_flavor takes "igraph", the default, or
"leidenalg", the backend scib itself uses. Both stored sources were
scored with leidenalg. The re-run tables were scored by evaluate from
multibench 0.2.1.
iASW and iF1 score every cell type, whereas scib's default scores only types confined to few batches. The re-run tables use the same rule.
metrics.attrs records how the scores were computed: the Leiden
backend, where the clusters came from, and the multibench and scib
versions. mtb.to_long keeps the first three in its scored_with
column, such as leidenalg/sweep/0.3.2.
Next: draw it next to the stored tables in Load and plot results.