Skip to content

Installation

pip install multibench-sc installs the API on Python 3.9 or newer. On macOS and Windows you can use the stored results, check your data with scan, score embeddings and plot. Running a method needs Linux. Windows is not tested.

Summary

conda create -n multibench python=3.11 -y && conda activate multibench
pip install multibench-sc
python -c "import multibench as mtb; print(mtb.__version__)"

Step 1: Create an isolated environment

Any fresh Python 3.9+ environment works. Conda is optional: the method environments install without it.

terminal
conda create -n multibench python=3.11 -y
conda activate multibench
terminal
mamba create -n multibench python=3.11 -y
mamba activate multibench
terminal
python3.11 -m venv .venv
source .venv/bin/activate

On Windows, activate it with .venv\Scripts\activate.


Step 2: Install multibench

terminal
pip install multibench-sc
Dependencies

The install brings numpy, pandas, h5py, anndata, matplotlib, pyyaml, scanpy and scib. PyTorch, R and the method packages are not installed here. Each method's environment holds them.

Developer install

Clone the repository only to edit the notebooks or run the test suite:

terminal
git clone https://github.com/DSichang/scMultiBench.git
cd scMultiBench
pip install -e ".[dev]" --config-settings editable_mode=compat

Step 3: Verify the install

terminal
python -c "import multibench as mtb; print(mtb.__version__)"

These calls need no method environment:

import multibench as mtb

print(mtb.list_tasks())
print(mtb.list_methods(category="vertical"))
print(mtb.describe_layout("vertical"))   # how to lay out your own dataset

Method environments

Each method runs in its own Linux environment. Install only the ones you need. --packed downloads prebuilt archives. No conda binary is needed for the packed path.

terminal
# installed or missing, with the install line
multibench env doctor
# the environments a category needs, with sizes
multibench env plan --category vertical
# one method
multibench env install --methods SCALEX --packed --run
# one category
multibench env install --category vertical --packed --run
Details
  • mtb.config.DEFAULT.envs_dir is where they go: $MULTIBENCH_ENVS_DIR when set, else the envs directory of the conda or mamba on PATH, else ~/.cache/multibench/envs. multibench config shows the folder in use.
  • Without --packed, each environment is built from its lockfile. This needs conda or mamba and takes longer. A failed archive download also falls back to this build.
  • --flavor cpu|gpu picks the build. By default, a host without an NVIDIA GPU gets the smaller CPU build of the PyTorch environments that have one.
  • env plan totals: categories share environments, so two categories cost less than the sum of their totals. ? disk means the size on disk is not recorded.
  • env install skips environments that already exist, so a rerun resumes an interrupted install. With no --methods or --category, it installs all 26. Check the total with multibench env plan first.
  • mtb.env.status, plan, install, doctor and recipe mirror the CLI. mtb.env.install is a dry run until dry_run=False (reference).
Details: method scripts

The first mtb.run clones the method scripts from PYangLab/scMultiBench with git. They go to mtb.config.DEFAULT.repo_path, which is ~/.cache/multibench/scMultiBench_ref after pip install. multibench fetch --scripts fetches them ahead of time and prints the commit. --ref, or the variable MULTIBENCH_SCRIPTS_REF, fetches a given commit or tag.

On a host without network, fetch the method scripts on another machine and copy them. Then set MULTIBENCH_REPO_PATH to the copy, or pass repo_path=.

Google Colab

The tutorials open in Colab from their badge. Choose a GPU runtime first: Runtime -> Change runtime type -> T4 GPU. Each notebook then installs the environments of the methods it runs and runs them, with no conda needed.

Details

On a CPU runtime, an environment that has a smaller CPU build gets that build, and training methods run slower.

The environments each tutorial downloads, with a GPU and without one:

  • vertical 5.1 GB, 1.1 GB: Matilda and sciPENN
  • diagonal 0.9 GB: iNMF and online_iNMF
  • mosaic 5.5 GB, 1.8 GB: StabMap and scMoMaT
  • cross 3.0 GB, 1.4 GB: StabMap and sciPENN
  • end-to-end 3.0 GB, 0.6 GB: Matilda

Get the data

Each tutorial downloads its dataset on first run. Datasets go to mtb.config.DEFAULT.data_path: ~/.cache/multibench/data after pip install, <repo>/data in a developer install.

import multibench as mtb

mtb.data.fetchable()      # ['D11', 'D28', 'D45', 'D46', 'D52']
mtb.data.fetch("D11")     # no-op when already present
Details

The demo datasets are D11 (vertical, 11 MB), D28 (diagonal, 137 MB), D45 (mosaic, 290 MB), D46 (mosaic, 97 MB) and D52 (cross, 179 MB).

MULTIBENCH_DATA_PATH sets the data path. Without it, and with XDG_CACHE_HOME set, the data path after pip install is $XDG_CACHE_HOME/multibench/data. multibench config shows the path in use.

To download a dataset by hand, for example D11:

terminal
DATA=$(multibench config --get data_path)
URL=https://github.com/DSichang/scMultiBench/releases/download/data-v1
mkdir -p "$DATA"
wget -qO- "$URL/D11.tar.gz" | tar xz -C "$DATA"

The full collection of processed datasets is linked from the scMultiBench README. Unpack a dataset under the data path, or anywhere else and pass data_path= (--data-path on the command line). dataset is the folder name, and data_path is the folder that contains it.


On a cluster with offline compute nodes

Prepare everything on a login node with network. The compute nodes then need no download. Keep the folders below on storage that the compute nodes can read.

login node
python3 -m venv /shared/multibench/venv
source /shared/multibench/venv/bin/activate
# 0.3.2 or newer: fetch --scripts and --assume-gpu need it
pip install "multibench-sc>=0.3.2"
export MULTIBENCH_ENVS_DIR=/shared/multibench/envs
export MULTIBENCH_DATA_PATH=/shared/multibench/data
export MULTIBENCH_REPO_PATH=/shared/multibench/scripts
# export MULTIBENCH_SCRIPTS_REF=<commit>  # optional: pin the scripts
# --flavor gpu: for jobs on GPU nodes
multibench env install --methods StabMap,scMoMaT --packed --flavor gpu --run
multibench fetch --scripts
# optional: a demo mosaic dataset to test the setup
multibench fetch D46
# LAB: your dataset folder under $MULTIBENCH_DATA_PATH
# --assume-gpu: check GPU-only methods for the GPU nodes, not this node
# --strict: exit 1 when any method in --methods has no runnable row
multibench scan LAB --category mosaic --methods StabMap,scMoMaT \
    --strict --assume-gpu

Submit one job per method, each with its own --out-dir. Request a GPU only for a method that uses one. multibench info M shows whether a method uses a GPU.

job.sh (Slurm)
#!/bin/bash
#SBATCH --time=05:00:00
source /shared/multibench/venv/bin/activate
export MULTIBENCH_ENVS_DIR=/shared/multibench/envs
export MULTIBENCH_DATA_PATH=/shared/multibench/data
export MULTIBENCH_REPO_PATH=/shared/multibench/scripts
# export MULTIBENCH_SCRIPTS_REF=<commit>  # optional: pin the scripts
# exits 3 when the method fails, so Slurm marks the job as failed
multibench run-all LAB --category mosaic --methods "$1" \
    --out-dir "runs/$1" --timeout 14400
login node
# StabMap does not use a GPU; scMoMaT uses one when present
sbatch job.sh StabMap
sbatch --gres=gpu:1 job.sh scMoMaT

When the jobs have finished, plot them together:

login node
multibench plot bubble --input runs/StabMap --input runs/scMoMaT --out lab.pdf
Details

env plan gives download sizes. An unpacked environment is larger. Check it with du after the first install.

Each input file is read as a dense matrix. Request at least features x cells x 8 bytes of memory per input file.

The runtimes that multibench info M prints were observed on a GPU host. For a method that uses a GPU, allow several times longer on a CPU node.


Upgrading from 0.2.1

0.3 renamed a few functions and merged the metric selectors into one metrics= argument. Some old spellings still work in 0.3 with a DeprecationWarning. Others were removed and raise an error. The old -> new table is on the Changes page.


Troubleshooting

OSError: Environment <env> of <method> is not installed

The method's environment is missing. Run the install command the message names, for example multibench env install --methods <method> --packed --run. scan() shows the same per method in its env_ok column.

OSError: Methods run only on Linux ...

Methods cannot run on macOS or Windows. Run the call on a Linux machine. dry_run=True (--dry-run) previews the command on this computer.

FileNotFoundError: ... 'conda' from mtb.run

No method environment is installed yet and this host has no conda. Install the method's environment with multibench env install --methods <method> --packed --run.

OSError naming envs_dir, with MULTIBENCH_RUN_MODE=prefix

Prefix mode is forced, but the environment is not under envs_dir. Install it (multibench env install --methods <method> --packed --run), point MULTIBENCH_ENVS_DIR at the folder that holds it, or unset MULTIBENCH_RUN_MODE.

RuntimeError: Conda is not installed on this computer. Environment <env> has a packed archive.

Only a lockfile build needs conda. Pass packed=True (--packed on the command line).

ImportError: No module named scib from mtb.evaluate

multibench was installed with --no-deps or into another interpreter. Run pip install multibench-sc in the environment you import from.

cLISI / iLISI are NaN, with a warning naming knn_graph.o

Both need scib's compiled LISI helper, which ships as a Linux x86-64 binary. On other platforms evaluate() rebuilds it automatically when a C++ compiler (g++, c++ or clang++) is on PATH.

Without a compiler, install one (Xcode command-line tools on macOS) and rerun evaluate in a new Python session, or run the g++ command the warning prints. The other metrics are computed either way.

unsupported input format from mtb.io.to_canonical

Accepted inputs: .h5ad, .h5mu, .csv / .tsv, .loom, in-memory AnnData / MuData, or a .h5 file already in the package's layout. .loom needs pip install "multibench-sc[loom]". Convert other formats first. For a whole dataset folder, use mtb.io.export_dataset or multibench convert.

If the problem remains, open an issue with the output of python -c "import multibench as mtb; print(mtb.__version__)" and pip freeze.


Next: the Quickstart, or run a method with the Run a method tutorial.