Skip to content

Discover methods

These calls show which of the 36 methods fit your data and your task. They need no environments, GPU or data.

Summary

import multibench as mtb

mtb.list_tasks()                        # the tasks methods are scored on
mtb.list_methods(category="vertical")   # every method in a category
# filter by what your data has
mtb.find_methods(category="diagonal", modalities=["rna", "atac"],
                 needs_labels=False)
mtb.method_info("SCALEX")               # everything known about one method
# rank by the stored scores
mtb.recommend("vertical", modalities=["rna", "adt"])
# the benchmark and the method, one line each
print(mtb.cite(["SCALEX"]))

The four integration categories

Pick the category that matches your experimental design and pass it as category= first. A method written for one design cannot take data from another.

Category Design Typical modalities
vertical Several modalities measured in the same cells RNA+ADT, RNA+ATAC, RNA+ADT+ATAC
diagonal Modalities measured in different cells, bridged by shared features RNA batch + ATAC batch
mosaic Several batches, only some sharing a modality, joined by a paired batch mixed, partially paired
cross Several batches, each with every modality RNA+ADT (the only cross data this package runs)

Several samples of 10x Multiome (RNA+ATAC) belong in vertical: one folder with all cells. To score batch mixing, pass each cell's sample as batch= to run_all or evaluate.


List what exists

list_methods filters by category only. To filter by task, use find_methods.

import multibench as mtb

mtb.list_tasks()
# ['batch', 'clustering', 'dimension_reduction']

mtb.list_methods()                       # all 36 methods
mtb.list_methods(category="vertical")    # paired-design methods only
mtb.find_methods("diagonal", task="clustering")

Filter by what your data has

find_methods filters by task and by what your data has: the modalities you measured, how your ATAC is encoded, and whether you have cell-type labels.

Each method reads ATAC either as peaks or as gene activity. Give each method the representation it reads: atac= keeps the methods that read yours.

# unpaired RNA and ATAC batches, no cell-type labels
mtb.find_methods(category="diagonal", modalities=["rna", "atac"],
                 needs_labels=False)
# ['SCALEX', 'Seurat_v5', 'VIPCCA', 'Portal', ...]

mtb.find_methods(category="diagonal", atac="peak")   # ATAC as peaks
mtb.find_methods(category="vertical", tunable=True)  # settable parameters
Details

A method matches when one of its variants meets all the filters together. A variant is the version of a method for one category and one set of inputs. None, the default, turns a filter off.

modalities= keeps the variants that take every listed type, so ["rna"] is broad and ["rna", "atac"] stricter. "atac" matches any ATAC input, and protein means adt.

The other ATAC tokens select by what the method reads. "gas", "gene_activity" and "atac_gas" keep the methods that read gene activity. "peak" and "atac_peak" keep those that read peaks, including moETM, scMM and iPOLNG, whose input is named atac_gas. Naming both representations keeps MultiMAP and Seurat_v3, which read both files.

atac= takes "peak" or "gene_activity", and the aliases "peaks" and "gas". A representation token that contradicts atac= raises ValueError.

mtb.scan and run_all use the same tokens. They keep a row only when its modalities are exactly the listed ones, so scMoMaT's rna+adt+atac row is not kept for ["rna", "atac"]. A representation token keeps every row whose method reads it, so "atac_peak" also keeps MultiMAP and Seurat_v3.

needs_labels is judged per variant. method_info(m)["needs_labels"] is True when any variant needs labels: for scMoMaT only its mosaic variant does. tunable=True keeps the methods whose command line exposes at least one hyperparameter for run(params=...).

An unknown token raises ValueError listing the valid values. A bare string such as modalities="rna" raises TypeError.

My ATAC is peaks only

Most diagonal methods read gene activity. With a peak matrix alone, check which methods can run:

  • GLUE runs on peaks. It reads peak names such as chr1:100-200, and mtb.run rewrites chr1_100_200 or chr1-100-200 into that form. It also needs the GENCODE annotation file. See Run a method.
  • Seurat_v5 needs a third dataset: a paired bridge with RNA and ATAC from the same cells. The package passes your RNA and ATAC files as that bridge, so it cannot run on unpaired data here.
  • MultiMAP and Seurat_v3 also need gene activity for the same ATAC cells.
  • The other ten need gene activity instead of peaks, for the ATAC cells of atac_peak.h5. Make it outside the package, for example with Signac GeneActivity. Save it with mtb.io.to_canonical(gas, "data/X/", modality="gas"). From R, write.csv(t(as.matrix(gas)), "gas.csv") writes the matrix as cells x genes. Pass "gas.csv" as gas.

mtb.scan then shows which methods have their files.

Details

mtb.scan and inputs_for(..., check=True) compare the cells of atac_gas.h5 with those of atac_peak.h5. Other cells, or the same cells in another order, fail the file check. Barcodes that differ only in a suffix such as -1 or -2 count as the same cell.

to_canonical writes the rows in the cell order of atac_peak.h5, which atac_cty.csv follows. The call warns when the gene-activity matrix holds other cells than atac_peak.h5. To re-order an atac_gas.h5 you already wrote, pass mtb.io.read_canonical(path) to to_canonical as the source.


Inspect a single method

method_info returns what the package knows about one method.

info = mtb.method_info("SCALEX")
info["env"], info["atac"], info["needs_labels"]
# ('scmb_torch', 'gene_activity', False)
info["supports"]    # one entry per variant: category, modalities, labels, ...
info["runtime"]     # {..., 'tier': 'slow', 'worst_sec': 2233, ...}
print(mtb.cite("SCALEX"))   # the benchmark and the method, one line each
Details

runtime["observed"] lists measured runs, with dataset, cells and seconds, and tier summarises them from fast to very_slow. Use it to set run_all(timeout=...). On a CPU, training methods take much longer than these times.

info["setup_hint"] names what a method needs before its first run, such as GLUE's annotation file. setup_hint is empty for most methods. mtb.params_for(method) lists what run(params=...) can change and what the script fixes.

deep_learning and output are in mtb.catalog.methods(), not in method_info.

mtb.cite("SCALEX", "GLUE") takes one id per argument or one list, and fmt="bibtex" returns .bib entries. Unknown ids raise KeyError with a did-you-mean hint.


Rank from stored results

recommend ranks a category's methods by their stored scores.

mtb.recommend("vertical", modalities=["rna", "adt"])
#   method  grand_score  n_datasets  ...  output_kind  datasets
mtb.recommend("diagonal", atac="peak")   # only methods that read peaks

grand_score compares a category's methods with each other. On each dataset, the best of the category's scored methods gets 1 and the lowest gets 0. grand_score is the mean over the datasets a method was scored on. It is not a quality score.

The stored scores cover few datasets: for vertical, only D11. Read the ranking as a hint. The full benchmark ranking is in the interactive explorer and in Fig. 6 of the paper. Methods with needs_labels=True train on your cell types, so compare them only with each other.

Details

coverage is the share of the ranked datasets a method was scored on, and datasets names them.

The Seurat_v5 script reads its bridge as two extra files. Its stored diagonal scores come from runs with a separate paired bridge dataset. They do not mean that it runs on your unpaired folder.

source="rerun" ranks the package's re-run tables instead of the published ones. The two hold different methods: for vertical D11 the published table has six and the re-run table 13. So the two can name different winners. To rank your own method alongside, score it with the leidenalg backend and pass long_df=pd.concat([mtb.load_results(category), mine]).

Methods that fit the category but have no stored scores are listed last, with grand_score NaN and coverage 0.0. One UserWarning names them and any datasets left out for having too few methods. The same warning says when modalities= or atac= names a modality that no ranked dataset measured.

modalities= and atac= follow the find_methods rule above. They only shorten the list, so a filtered list can start below 1 and end above 0.

metrics= picks the metric family. The default ranks on the clustering metrics. needs_labels follows the category's variants: scMoMaT is False in a vertical ranking.


The full benchmark catalog

mtb.catalog returns the benchmark's method, dataset and metric tables as DataFrames, and maps free-text names to canonical ones.

from multibench import catalog

catalog.methods()                  # one row per method: language, atac, ...
catalog.datasets()                 # one row per dataset, with its category
catalog.metrics()                  # what each metric measures, range, direction
catalog.canonical_id("Seurat v5")  # 'Seurat_v5'
catalog.canonical_metric("kbet")   # 'kBET'
Details

catalog.datasets() also lists the ids that exist only as stored results, such as the re-run subsamples D11s and D28s. The simulated column marks the simulated datasets.

canonical_id returns an unknown name with its spaces and dots replaced by _. canonical_id(name, strict=True) raises KeyError instead. catalog.known_metrics() lists the ten scIB codes and PCR.


Before running

Check your folder with mtb.scan("MYDATA", category). It reports files_ok and env_ok per method. Then run them with run_all.


Next: run a chosen method in Run a method, score it in Evaluate a run, and compare methods in Load and plot results.