Skip to content

Configuration (mtb.config)

The paths and settings every function reads. Change a field of mtb.config.DEFAULT to use your own data folder, environment folder or Leiden backend.

Name Summary
mtb.config.Config Resolved filesystem paths; override fields to point at custom locations.
mtb.AmbiguousVariantError Raised when several variants of a method fit and the call must pick one.

Environment variables

variable effect
MULTIBENCH_DATA_PATH The folder that holds the dataset folders (Config.data_path).
MULTIBENCH_ENVS_DIR Where method environments are installed and looked up (Config.envs_dir).
MULTIBENCH_REPO_PATH Where the method scripts are, or are downloaded to on first use (Config.repo_path).
MULTIBENCH_SCRIPTS_REF A commit or tag of the method scripts to download instead of the default branch.
MULTIBENCH_RUN_MODE prefix or conda: how mtb.run enters an environment. How the environment is entered describes both.
MULTIBENCH_DEBUG 1 makes the multibench CLI print the full traceback of an error.

A value assigned in Python wins over the variable. multibench config prints each path and where it came from.

mtb.config.Config dataclass

Config

Resolved filesystem paths; override fields to point at custom locations.

mtb.config.DEFAULT is the instance every function reads. Set its fields directly, or pass a path explicitly where a function takes data_path= / result_path=.

ATTRIBUTES DESCRIPTION
result_path

Result tables shipped with the package, read by mtb.load_results. Default <package root>/multibench/result.

type Path

files_path

Catalog CSVs read by mtb.catalog (method.csv, dataset.csv, metric_full.csv). Default <package root>/multibench/files.

type Path

repo_path

Checkout holding the upstream tools_scripts/ (the method scripts), cloned on first use when absent or empty. Default $MULTIBENCH_REPO_PATH, else <base>/scMultiBench_ref.

type Path

data_path

Data root that holds the dataset folders; mtb.data.fetch, mtb.scan and mtb.run_all use it. Default $MULTIBENCH_DATA_PATH, else <base>/data.

type Path

leiden_flavor

Leiden backend of the scIB clustering sweep in mtb.evaluate: "igraph" (default, faster) or "leidenalg" (the backend of both stored sources).

type str

envs_dir

Where the method environment prefixes live (<envs_dir>/<env>); resolved on first read (order in Notes).

type Path

Examples

>>> import multibench as mtb
>>> mtb.config.DEFAULT.data_path = "/scratch/data"
>>> # before mtb.env.install(...)
>>> mtb.config.DEFAULT.envs_dir = "/scratch/envs"
>>> # the backend of both stored sources, not igraph
>>> mtb.config.DEFAULT.leiden_flavor = "leidenalg"
>>> cfg = mtb.config.Config(data_path="/data/mine")  # a separate instance
Notes

Environment variables. Set them in the shell, a job script or a module file, and every process that sees them uses the paths. A value assigned in Python wins over the variable. multibench config prints each resolved path and where it came from. The variables and the fields they set:

MULTIBENCH_DATA_PATH   data_path
MULTIBENCH_REPO_PATH   repo_path
MULTIBENCH_ENVS_DIR    envs_dir

Where <base> is. The repository root in a checkout or editable install (pyproject.toml next to the package). For a wheel install it is the per-user cache ~/.cache/multibench ($XDG_CACHE_HOME honoured).

Assigning paths. data_path, repo_path and envs_dir accept a string and store a pathlib.Path; assigning None returns the field to its default. Assign result_path and files_path as pathlib.Path objects: mtb.load_results and mtb.catalog use them as they are.

How envs_dir is resolved. Lazily, on first read, from the first of:

  1. the MULTIBENCH_ENVS_DIR environment variable;
  2. the first writable envs directory of the conda/mamba found on PATH;
  3. ~/.cache/multibench/envs ($XDG_CACHE_HOME honoured).

It is what mtb.env.install unpacks packed archives into and what the runner's prefix mode activates. The first read may run conda info --json (once per process, not at import); assigning a value skips the probe.

Leiden backends. "igraph" is scanpy's igraph implementation, several times faster; "leidenalg" is the backend both stored sources (published and re-run) were computed with.

Method scripts. The first run fetches them from GitHub into repo_path (multibench fetch --scripts does it ahead). Set MULTIBENCH_SCRIPTS_REF to a commit or tag to fetch that version instead of the default branch; scripts already present must be at that ref, or runs refuse and mtb.scan blocks every row. multibench config and every run record show the commit in use (scripts_commit).

See Also

mtb.env.install : provisions the method envs under envs_dir. mtb.data.fetch : downloads reference datasets into data_path.

mtb.AmbiguousVariantError

Bases: ValueError, KeyError

Raised when several variants of a method fit and the call must pick one.

Raised by mtb.params_for and mtb.inputs_for; the message spells out the call that selects one variant.

Examples

>>> from multibench import params_for, AmbiguousVariantError
>>> params_for("Matilda")  # raises: rna+adt or rna+atac
>>> params_for("Matilda", "vertical", ["rna", "adt"])  # selects one
Notes

When it fires. category= / modalities= leave more than one variant - e.g. Matilda has an rna+adt and an rna+atac vertical variant, so params_for("Matilda") cannot pick:

  • mtb.inputs_for (without modalities=) and mtb.params_for(dataset=) (without category= or modalities=) first let the dataset folder decide, and raise only when it settles nothing;
  • mtb.labels_for never raises it; it falls back to the default label order.

What to do. Pass modalities= (and category=) exactly as the message shows.

Catching it. It is a ValueError and also a KeyError, so except KeyError catches it too. The package uses KeyError for unknown ids, such as a mistyped method name. str(exc) is the plain message, without KeyError quoting.

See Also

mtb.method_info : supports lists every variant with its category and modalities