Installation¶
pip install multibench-sc installs the API on Python 3.9 or newer. On macOS
and Windows you can use the stored results, check your data with scan,
score embeddings and plot. Running a method needs Linux. Windows is not
tested.
Summary
Step 1: Create an isolated environment¶
Any fresh Python 3.9+ environment works. Conda is optional: the method environments install without it.
Step 2: Install multibench¶
Dependencies
The install brings numpy, pandas, h5py, anndata, matplotlib,
pyyaml, scanpy and scib. PyTorch, R and the method packages are not
installed here. Each method's environment holds them.
Developer install
Clone the repository only to edit the notebooks or run the test suite:
Step 3: Verify the install¶
These calls need no method environment:
import multibench as mtb
print(mtb.list_tasks())
print(mtb.list_methods(category="vertical"))
print(mtb.describe_layout("vertical")) # how to lay out your own dataset
Method environments¶
Each method runs in its own Linux environment. Install only the ones you
need. --packed downloads prebuilt archives.
No conda binary is needed for the packed path.
# installed or missing, with the install line
multibench env doctor
# the environments a category needs, with sizes
multibench env plan --category vertical
# one method
multibench env install --methods SCALEX --packed --run
# one category
multibench env install --category vertical --packed --run
Details
mtb.config.DEFAULT.envs_diris where they go:$MULTIBENCH_ENVS_DIRwhen set, else the envs directory of the conda or mamba onPATH, else~/.cache/multibench/envs.multibench configshows the folder in use.- Without
--packed, each environment is built from its lockfile. This needs conda or mamba and takes longer. A failed archive download also falls back to this build. --flavor cpu|gpupicks the build. By default, a host without an NVIDIA GPU gets the smaller CPU build of the PyTorch environments that have one.env plantotals: categories share environments, so two categories cost less than the sum of their totals.? diskmeans the size on disk is not recorded.env installskips environments that already exist, so a rerun resumes an interrupted install. With no--methodsor--category, it installs all 26. Check the total withmultibench env planfirst.mtb.env.status,plan,install,doctorandrecipemirror the CLI.mtb.env.installis a dry run untildry_run=False(reference).
Details: method scripts
The first mtb.run clones the method scripts from PYangLab/scMultiBench
with git. They go to mtb.config.DEFAULT.repo_path, which is
~/.cache/multibench/scMultiBench_ref after pip install.
multibench fetch --scripts fetches them ahead of time and prints the
commit. --ref, or the variable MULTIBENCH_SCRIPTS_REF, fetches a
given commit or tag.
On a host without network, fetch the method scripts on another machine
and copy them. Then set MULTIBENCH_REPO_PATH to the copy, or pass
repo_path=.
Google Colab¶
The tutorials open in Colab from their badge. Choose a GPU runtime first: Runtime -> Change runtime type -> T4 GPU. Each notebook then installs the environments of the methods it runs and runs them, with no conda needed.
Details
On a CPU runtime, an environment that has a smaller CPU build gets that build, and training methods run slower.
The environments each tutorial downloads, with a GPU and without one:
- vertical 5.1 GB, 1.1 GB: Matilda and sciPENN
- diagonal 0.9 GB: iNMF and online_iNMF
- mosaic 5.5 GB, 1.8 GB: StabMap and scMoMaT
- cross 3.0 GB, 1.4 GB: StabMap and sciPENN
- end-to-end 3.0 GB, 0.6 GB: Matilda
Get the data¶
Each tutorial downloads its dataset on first run. Datasets go to
mtb.config.DEFAULT.data_path: ~/.cache/multibench/data after
pip install, <repo>/data in a developer install.
import multibench as mtb
mtb.data.fetchable() # ['D11', 'D28', 'D45', 'D46', 'D52']
mtb.data.fetch("D11") # no-op when already present
Details
The demo datasets are D11 (vertical, 11 MB), D28 (diagonal, 137 MB),
D45 (mosaic, 290 MB), D46 (mosaic, 97 MB) and D52 (cross, 179 MB).
MULTIBENCH_DATA_PATH sets the data path. Without it, and with
XDG_CACHE_HOME set, the data path after pip install is
$XDG_CACHE_HOME/multibench/data. multibench config shows the path in
use.
To download a dataset by hand, for example D11:
DATA=$(multibench config --get data_path)
URL=https://github.com/DSichang/scMultiBench/releases/download/data-v1
mkdir -p "$DATA"
wget -qO- "$URL/D11.tar.gz" | tar xz -C "$DATA"
The full collection of processed datasets is linked from the
scMultiBench README.
Unpack a dataset under the data path, or anywhere else and pass
data_path= (--data-path on the command line). dataset is the folder
name, and data_path is the folder that contains it.
On a cluster with offline compute nodes¶
Prepare everything on a login node with network. The compute nodes then need no download. Keep the folders below on storage that the compute nodes can read.
python3 -m venv /shared/multibench/venv
source /shared/multibench/venv/bin/activate
# 0.3.2 or newer: fetch --scripts and --assume-gpu need it
pip install "multibench-sc>=0.3.2"
export MULTIBENCH_ENVS_DIR=/shared/multibench/envs
export MULTIBENCH_DATA_PATH=/shared/multibench/data
export MULTIBENCH_REPO_PATH=/shared/multibench/scripts
# export MULTIBENCH_SCRIPTS_REF=<commit> # optional: pin the scripts
# --flavor gpu: for jobs on GPU nodes
multibench env install --methods StabMap,scMoMaT --packed --flavor gpu --run
multibench fetch --scripts
# optional: a demo mosaic dataset to test the setup
multibench fetch D46
# LAB: your dataset folder under $MULTIBENCH_DATA_PATH
# --assume-gpu: check GPU-only methods for the GPU nodes, not this node
# --strict: exit 1 when any method in --methods has no runnable row
multibench scan LAB --category mosaic --methods StabMap,scMoMaT \
--strict --assume-gpu
Submit one job per method, each with its own --out-dir. Request a GPU only
for a method that uses one. multibench info M shows whether a method uses a
GPU.
#!/bin/bash
#SBATCH --time=05:00:00
source /shared/multibench/venv/bin/activate
export MULTIBENCH_ENVS_DIR=/shared/multibench/envs
export MULTIBENCH_DATA_PATH=/shared/multibench/data
export MULTIBENCH_REPO_PATH=/shared/multibench/scripts
# export MULTIBENCH_SCRIPTS_REF=<commit> # optional: pin the scripts
# exits 3 when the method fails, so Slurm marks the job as failed
multibench run-all LAB --category mosaic --methods "$1" \
--out-dir "runs/$1" --timeout 14400
# StabMap does not use a GPU; scMoMaT uses one when present
sbatch job.sh StabMap
sbatch --gres=gpu:1 job.sh scMoMaT
When the jobs have finished, plot them together:
Details
env plan gives download sizes. An unpacked environment is larger.
Check it with du after the first install.
Each input file is read as a dense matrix. Request at least features x cells x 8 bytes of memory per input file.
The runtimes that multibench info M prints were observed on a GPU
host. For a method that uses a GPU, allow several times longer on a CPU
node.
Upgrading from 0.2.1¶
0.3 renamed a few functions and merged the metric selectors into one
metrics= argument. Some old spellings still work in 0.3 with a
DeprecationWarning. Others were removed and raise an error. The old -> new
table is on the Changes page.
Troubleshooting¶
OSError: Environment <env> of <method> is not installed
The method's environment is missing. Run the install command the message
names, for example multibench env install --methods <method> --packed --run.
scan() shows the same per method in its env_ok column.
OSError: Methods run only on Linux ...
Methods cannot run on macOS or Windows. Run the call on a Linux machine.
dry_run=True (--dry-run) previews the command on this computer.
FileNotFoundError: ... 'conda' from mtb.run
No method environment is installed yet and this host has no conda.
Install the method's environment with
multibench env install --methods <method> --packed --run.
OSError naming envs_dir, with MULTIBENCH_RUN_MODE=prefix
Prefix mode is forced, but the environment is not under envs_dir.
Install it (multibench env install --methods <method> --packed --run),
point MULTIBENCH_ENVS_DIR at the folder that holds it, or unset
MULTIBENCH_RUN_MODE.
RuntimeError: Conda is not installed on this computer. Environment <env> has a packed archive.
Only a lockfile build needs conda. Pass packed=True (--packed on the
command line).
ImportError: No module named scib from mtb.evaluate
multibench was installed with --no-deps or into another interpreter.
Run pip install multibench-sc in the environment you import from.
cLISI / iLISI are NaN, with a warning naming knn_graph.o
Both need scib's compiled LISI helper, which ships as a Linux x86-64
binary. On other platforms evaluate() rebuilds it automatically when a
C++ compiler (g++, c++ or clang++) is on PATH.
Without a compiler, install one (Xcode command-line tools on macOS) and
rerun evaluate in a new Python session, or run the g++ command the
warning prints. The other metrics are computed either way.
unsupported input format from mtb.io.to_canonical
Accepted inputs: .h5ad, .h5mu, .csv / .tsv, .loom, in-memory
AnnData / MuData, or a .h5 file already in the package's layout.
.loom needs pip install "multibench-sc[loom]". Convert other formats
first. For a whole dataset folder, use mtb.io.export_dataset or
multibench convert.
If the problem remains,
open an issue with the
output of python -c "import multibench as mtb; print(mtb.__version__)" and
pip freeze.
Next: the Quickstart, or run a method with the Run a method tutorial.