Skip to content

Score

Score a run output with scIB metrics, and reshape the scores into the long table that mtb.load_results returns. To compare with the stored tables, set mtb.config.DEFAULT.leiden_flavor = "leidenalg" first. The API overview summarises the metrics= selector.

Name Summary
mtb.evaluate Compute scIB metrics for a run output (an embedding) against cell-type labels.
mtb.to_long Reshape mtb.evaluate's scores into the long table of mtb.load_results.

mtb.evaluate

evaluate(
    output,
    labels=None,
    *,
    category: str | None = None,
    batch=None,
    metrics=None,
    clustering=None,
    obsm: str = "X_emb",
    label_order=None,
    verbose: bool = True,
) -> DataFrame

Compute scIB metrics for a run output (an embedding) against cell-type labels.

Pass the embedding a run produced and one label per cell. Reshape the result with mtb.to_long to plot it or combine it with mtb.load_results.

PARAMETERS DESCRIPTION
output

Embedding, cells x dims: array, DataFrame, sparse matrix, AnnData or MuData (.obsm[obsm]; labels/batch may name .obs columns), or file path (Notes).

labels

Cell types, one per cell (required): CSV path(s), a mtb.labels_for dict, a 1-D array-like, or an .obs column name (forms in Notes).

default None

category

Integration category, one of mtb.list_categories(); validated, otherwise unused (the metrics do not depend on it). Lets a call mirror mtb.run.

type str default None

batch

Batch labels, one per cell, in the forms labels accepts; needed for the batch metrics unless labels lists two or more files.

default None

metrics

None = every applicable metric except kBET; or "clustering", "batch", "all", or a list of codes such as ["ARI", "NMI"].

type None, str or list of str default None

clustering

Precomputed cluster assignment, in the forms labels accepts or an .h5 path; None = derive one with the Leiden sweep.

default None

obsm

.obsm key of an AnnData, MuData or .h5ad output, such as 'X_pca'; 'X' means .X.

type str default 'X_emb'

label_order

Keys of a multi-entry labels dict, in the method's cell order; a subset selects those files. None works for an unchanged mtb.labels_for dict (Notes).

type list of str default None

verbose

True prints one stderr line when the Leiden sweep starts on more than 2,000 cells; False never prints.

type bool default True

RETURNS DESCRIPTION
DataFrame

One row per metric, indexed by the canonical metric name (ARI, NMI, ...), with one column Value. attrs records how the scores were computed (Notes).

RAISES DESCRIPTION
ValueError

Missing or misaligned labels, an unknown category or metric, or batch metrics without batch labels.

FileNotFoundError

An output or label path does not exist.

TypeError

An unsupported input type, or a retired keyword (Notes).

Examples

>>> import multibench as mtb
>>> labels = mtb.labels_for("D11")                    # {'cty': '.../cty.csv'}
>>> scores = mtb.evaluate(res.output, labels=labels)  # res = mtb.run(...)
>>> # ASW and cLISI need no Leiden sweep
>>> mtb.evaluate(res.output, labels=labels, metrics=["ASW", "cLISI"])
>>> mtb.evaluate(adata, labels="celltype", obsm="X_pca")
>>> mtb.evaluate(mdata, labels="celltype", batch="sample", obsm="X_joint",
...              metrics="all")
Notes

Label forms. labels may be:

  • a CSV path (header row; column x, the only column, or the last of two when the first is a barcode index);
  • a list of CSV paths, concatenated in that order (multi-batch datasets: [cty1, cty2, cty3]);
  • a {name: path} dict, such as mtb.labels_for returns (order rules under Label dicts.);
  • a 1-D ndarray / Series / Categorical / list, or a single-column DataFrame;
  • when output is an AnnData or MuData, the name of an .obs column.

A multi-column CSV / DataFrame raises: pass the one column (df["celltype"]) itself. batch and clustering take the same forms except a multi-entry dict; for clustering, a path that is not .csv / .tsv / .txt is read as an h5 from /obs/cluster_leiden.

Label dicts. A dict with several entries needs a known order:

  • a dict from mtb.labels_for, unchanged, is used as it is, in the order labels_for gave it (with category and method, that method's cell order);
  • any other dict goes in as is only in the default order (cty1, cty2, ... numerically; rna before adt before atac) and otherwise needs label_order= naming its keys;
  • label_order=list(d) trusts the dict's own order. A one-entry dict needs no label_order.

Metric selection. None computes the clustering family (ARI, NMI, ASW, iASW, iF1, cLISI), plus ASW_batch, GC, iLISI when a batch vector is available (batch= or a list of label files). kBET is never included by default: it shells out to R and takes hours on large datasets.

  • "clustering" / "batch" - that family; "all" - both ("batch" and "all" need the batch vector or raise);
  • a list of codes - exactly those (case/alias tolerant, ["ari"] works); the Leiden sweep runs only when one of them needs it (ARI, NMI, iF1), and "kBET" in the list turns kBET on.

Valid codes: mtb.plot.CLUSTERING_METRICS + mtb.plot.BATCH_METRICS. An unknown code raises ValueError listing the valid ones; a bare code string (metrics="ARI") raises and points at the list form.

Cost and Leiden backend. ARI, NMI and iF1 need the scIB optimal-resolution Leiden sweep (10 resolutions on a kNN graph of the embedding). mtb.config.DEFAULT.leiden_flavor sets its backend: "igraph" (default) or "leidenalg", the classic backend scib itself runs. With leidenalg the sweep takes tens of seconds for a few thousand cells and minutes for ~10^4; igraph is several times faster.

To skip it, name only metrics that do not need it in metrics=[...] (ASW, iASW, cLISI, the batch family). clustering= removes the need for ARI / NMI only; iF1 always sweeps.

Metric definitions. All metrics are scaled so that higher is better. Most lie between 0 and 1; ARI can be slightly below 0, which means a random clustering.

Each metric is a scib function on the embedding. The package requires scib >= 1.1; attrs["scib_version"] records the installed version.

The metrics look at different neighbourhoods:

  • the Leiden sweep (ARI, NMI, iF1) and GC use scanpy's default neighbour graph (15 neighbours);
  • cLISI and iLISI: scib builds scanpy's 15-neighbour graph from the embedding. For each cell it takes the 90 cells with the shortest paths on that graph, where the length of an edge is its connectivity weight (obsp["connectivities"]), not its distance. Perplexity 30. A cell with fewer than 90 reachable cells counts as one label;
  • ASW, iASW and ASW_batch use distances in the embedding;
  • kBET builds its own neighbour graph.

The sweep clusters at 10 resolutions (0.2 to 2.0) and keeps the one with the highest NMI. iso_threshold = number of batches + 1 counts every cell type as isolated; scib's default counts only types found in few batches.

The scib call behind each code:

code       scib function         arguments
ARI, NMI   ari, nmi              Leiden clusters, best sweep resolution
ASW        silhouette            cell-type labels; rescaled to 0-1 by scib
iASW       isolated_labels_asw   iso_threshold = number of batches + 1
iF1        isolated_labels_f1    same threshold; best F1 over the sweep
cLISI      clisi_graph           type_="embed"; 90 nearest by path over
                                 the connectivity weights of the
                                 15-neighbour graph; scaled to 0-1
ASW_batch  silhouette_batch      1 - |batch silhouette| per cell type
GC         graph_connectivity    on the 15-neighbour graph
iLISI      ilisi_graph           type_="embed"; 90 nearest by path over
                                 the connectivity weights of the
                                 15-neighbour graph; scaled to 0-1
kBET       kBET                  computed only when named in metrics=[...]

cLISI and iLISI take the median m of the per-cell scores and scale it: cLISI = (L - m)/(L - 1), iLISI = (m - 1)/(B - 1), with L cell types and B batches.

The re-run tables were scored by multibench 0.2.1's evaluate with these definitions and the leidenalg backend. The published tables were computed by the benchmark; see the paper's Methods.

Batch. The batch family (ASW_batch, GC, iLISI, kBET) needs batch labels. When labels is a list (or dict) of two or more files and batch is omitted, the file of origin (1, 2, ...) serves as the batch, the same rule mtb.run_all applies. batch= given with a metrics selection that has no batch metric changes nothing, and a UserWarning says so.

Cell order. Arrays, lists and label files without a cell-id column are matched by position to the rows of output. A Series / DataFrame with a non-default index is aligned by cell id when output carries ids (an AnnData, or a DataFrame with a non-default index): rows are reindexed to the output's order, and a missing or extra id raises ValueError naming the first ones.

A CSV whose first column holds the output's cell ids, as obs[["batch"]].to_csv(path) writes it, is aligned by that column in the same way, and so is a list of such files. A first column where only some values are cell ids raises ValueError.

A first column of text that holds none of the cell ids is matched by position, with a UserWarning. Numbers there, such as R's row numbers, are matched by position unless they are exactly the output's cell ids.

When output is a bare array there is nothing to align against: the Series, or a CSV with text in its first column, is matched positionally and a UserWarning says so.

Result shape. The frame has the metric.csv shape: index = metric, one column Value, never empty. mtb.to_long makes the long frame (lowercase value) that load_results returns.

Provenance. attrs records how the scores were computed:

  • leiden_flavor - the backend of the Leiden sweep; None when no sweep ran;
  • clustering - where the clusters ARI and NMI scored came from: "sweep" or "user"; None without ARI and NMI;
  • multibench_version and scib_version - the installed versions.

mtb.to_long writes the first three to a scored_with column.

Output formats. A file path may be .h5 (dataset data, the benchmark's embedding.h5), .h5ad (read as an AnnData), .npy, or .csv / .tsv. A dims x cells array is transposed against the label count, with a warning.

Errors. ValueError covers:

  • missing labels, or a batch metric / family requested without batch labels;
  • an unknown category, metrics token or code;
  • a count that differs from the embedding's cells ('labels has M entries for N cells in the embedding.'), and cell-id mismatches when aligning;
  • ambiguous label files, and a multi-entry labels dict (other than an unchanged labels_for one) out of the default order without label_order;
  • unknown, repeated or no keys in label_order;
  • an .h5 output without dataset data.

FileNotFoundError names the missing path and the working directory. TypeError covers unsupported input types (non-array labels, label_order with non-dict labels, ...).

Retired keywords. task=, family= and only= still work as spellings of metrics=, with a DeprecationWarning; slow_metrics=, column= and metric_set= raise TypeError naming the replacement. The Changes page lists them.

See Also

mtb.labels_for : the label files of a dataset, in the method's cell order.

mtb.to_long : reshapes the result into the long results frame.

mtb.run_all : runs and scores every runnable method on a dataset.

mtb.to_long

to_long(
    value_df,
    *,
    method: str,
    dataset: str | None = None,
    category: str | None = None,
    clustering: str = "default",
    source: str = "user",
) -> DataFrame

Reshape mtb.evaluate's scores into the long table of mtb.load_results.

The result has the columns of mtb.load_results and concatenates with the stored tables.

PARAMETERS DESCRIPTION
value_df

What mtb.evaluate returns (metrics as the index, one column Value), a Series indexed by metric, or that frame's CSV read back (forms in Notes).

type DataFrame or Series

method

Method id written into every row; your own name is fine.

type str

dataset

Dataset id written into every row; None writes "all".

type str default None

category

Integration category written into every row; None writes "user".

type str default None

clustering

Value of the clustering column.

type str default 'default'

source

Value of the source column; source="user" in mtb.load_results selects these rows again.

type str default 'user'

RETURNS DESCRIPTION
DataFrame

One row per metric, with the columns metric, value, method, dataset, category, clustering, source, scored_with. scored_with is "unknown" for a frame not from mtb.evaluate.

RAISES DESCRIPTION
ValueError

value_df is already long, lacks Value or string metric names, or repeats a name.

WARNS DESCRIPTION
UserWarning

value_df has no record of how it was scored. Its scored_with is "unknown".

Examples

>>> import multibench as mtb, pandas as pd
>>> wide = mtb.evaluate(emb, labels=labels, metrics=["ARI", "NMI"])
>>> mine = mtb.to_long(wide, method="MyMethod", dataset="D11",
...                    category="vertical")
>>> stored = mtb.load_results("vertical", dataset="D11", source="rerun")
>>> pd.concat([stored, mine]).to_csv("all.csv", index=False)
>>> mtb.load_results(result_path="all.csv", source="user")   # your rows only
Notes

Metric names. Names are canonicalised (ari -> ARI, kbet -> kBET); rows whose name is blank are dropped, and a name the package does not know is kept as written.

CSV read-back. pd.read_csv(path, index_col=0) reads wide.to_csv(path) back. A plain pd.read_csv(path) works too: its metric column Unnamed: 0 is taken as the names only when the columns are exactly Unnamed: 0, Value, it holds strings and the index none. So does a metric column (wide.to_csv(path, index_label="metric")). Metric names that are not strings, like row numbers 0, 1, ..., raise instead of being scored.

Column values. dataset=None writes "all". category=None writes "user", the value load_results gives a user file without a category column. The published tables use clustering values "louvain" / "kmeans" for their variants.

Provenance. The scored_with column of a frame from mtb.evaluate reads like "leidenalg/sweep/0.3.2": the Leiden backend, where the clusters came from (sweep or user) and the package version; none fills a part that did not apply. The same three values and scib_version are in attrs.

Any other frame, such as a hand-made one or a wide CSV read back, has no such record. Its rows get scored_with = "unknown" and a UserWarning. To keep the record, save the long table: mtb.to_long(...).to_csv(path, index=False).

The stored tables have no scored_with column, so it is NaN for their rows after pd.concat. load_results(result_path=...) keeps the column when it reads the CSV back.

Plot badges. The bubble figure shows ? for a name the package does not know, such as your own or a renamed re-run. To badge such a row as supervised, add a boolean needs_labels column.

Errors. All are ValueError:

  • an already long frame (columns metric, value, method) - pass it to the plot / load_results consumers directly;
  • no Value column - the message names the expected shape and, for a wide one-row frame, the df.T.set_axis(['Value'], axis=1) fix;
  • a metric name that is not a string (row numbers from a plain pd.read_csv of another shape) - the message names pd.read_csv(path, index_col=0);
  • every metric name blank;
  • two names that collapse onto one canonical metric (ari and ARI).
See Also

mtb.evaluate : computes the wide frame this function reshapes.

mtb.load_results : the stored scores, in the same long shape.

mtb.plot.bubble : plots a long frame.