Configuration (mtb.config)¶
The paths and settings every function reads. Change a field of
mtb.config.DEFAULT to use your own data folder, environment folder or Leiden
backend.
| Name | Summary |
|---|---|
mtb.config.Config |
Resolved filesystem paths; override fields to point at custom locations. |
mtb.AmbiguousVariantError |
Raised when several variants of a method fit and the call must pick one. |
Environment variables¶
| variable | effect |
|---|---|
MULTIBENCH_DATA_PATH |
The folder that holds the dataset folders (Config.data_path). |
MULTIBENCH_ENVS_DIR |
Where method environments are installed and looked up (Config.envs_dir). |
MULTIBENCH_REPO_PATH |
Where the method scripts are, or are downloaded to on first use (Config.repo_path). |
MULTIBENCH_SCRIPTS_REF |
A commit or tag of the method scripts to download instead of the default branch. |
MULTIBENCH_RUN_MODE |
prefix or conda: how mtb.run enters an environment. How the environment is entered describes both. |
MULTIBENCH_DEBUG |
1 makes the multibench CLI print the full traceback of an error. |
A value assigned in Python wins over the variable. multibench config prints
each path and where it came from.
mtb.config.Config
dataclass
¶
Resolved filesystem paths; override fields to point at custom locations.
mtb.config.DEFAULT is the instance every function reads. Set its
fields directly, or pass a path explicitly where a function takes
data_path= / result_path=.
| ATTRIBUTES | DESCRIPTION |
|---|---|
result_path |
Result tables shipped with the package, read by
type
|
files_path |
Catalog CSVs read by
type
|
repo_path |
Checkout holding the upstream
type
|
data_path |
Data root that holds the dataset folders;
type
|
leiden_flavor |
Leiden backend of the scIB clustering sweep in
type
|
envs_dir |
Where the method environment prefixes live (
type
|
Examples
>>> import multibench as mtb
>>> mtb.config.DEFAULT.data_path = "/scratch/data"
>>> # before mtb.env.install(...)
>>> mtb.config.DEFAULT.envs_dir = "/scratch/envs"
>>> # the backend of both stored sources, not igraph
>>> mtb.config.DEFAULT.leiden_flavor = "leidenalg"
>>> cfg = mtb.config.Config(data_path="/data/mine") # a separate instance
Notes
Environment variables. Set them in the shell, a job script or a
module file, and every process that sees them uses the paths. A value
assigned in Python wins over the variable. multibench config prints
each resolved path and where it came from. The variables and the fields
they set:
Where <base> is. The repository root in a checkout or editable
install (pyproject.toml next to the package). For a wheel install it
is the per-user cache ~/.cache/multibench ($XDG_CACHE_HOME
honoured).
Assigning paths. data_path, repo_path and envs_dir
accept a string and store a pathlib.Path; assigning None returns
the field to its default. Assign result_path and files_path as
pathlib.Path objects: mtb.load_results and mtb.catalog use
them as they are.
How envs_dir is resolved. Lazily, on first read, from the first
of:
- the
MULTIBENCH_ENVS_DIRenvironment variable; - the first writable envs directory of the conda/mamba found on PATH;
~/.cache/multibench/envs($XDG_CACHE_HOMEhonoured).
It is what mtb.env.install unpacks packed archives into and what the
runner's prefix mode activates. The first read may run conda info
--json (once per process, not at import); assigning a value skips
the probe.
Leiden backends. "igraph" is scanpy's igraph implementation,
several times faster; "leidenalg" is the backend both stored
sources (published and re-run) were computed with.
Method scripts. The first run fetches them from GitHub into
repo_path (multibench fetch --scripts does it ahead). Set
MULTIBENCH_SCRIPTS_REF to a commit or tag to fetch that version
instead of the default branch; scripts already present must be at that
ref, or runs refuse and mtb.scan blocks every row.
multibench config and every run record show the commit in use
(scripts_commit).
See Also
mtb.env.install : provisions the method envs under envs_dir.
mtb.data.fetch : downloads reference datasets into data_path.
mtb.AmbiguousVariantError
¶
Bases: ValueError, KeyError
Raised when several variants of a method fit and the call must pick one.
Raised by mtb.params_for and mtb.inputs_for; the message spells
out the call that selects one variant.
Examples
>>> from multibench import params_for, AmbiguousVariantError
>>> params_for("Matilda") # raises: rna+adt or rna+atac
>>> params_for("Matilda", "vertical", ["rna", "adt"]) # selects one
Notes
When it fires. category= / modalities= leave more than one
variant - e.g. Matilda has an rna+adt and an rna+atac vertical
variant, so params_for("Matilda") cannot pick:
mtb.inputs_for(withoutmodalities=) andmtb.params_for(dataset=)(withoutcategory=ormodalities=) first let the dataset folder decide, and raise only when it settles nothing;mtb.labels_fornever raises it; it falls back to the default label order.
What to do. Pass modalities= (and category=) exactly as the
message shows.
Catching it. It is a ValueError and also a KeyError, so
except KeyError catches it too. The package uses KeyError for
unknown ids, such as a mistyped method name. str(exc) is the plain
message, without KeyError quoting.
See Also
mtb.method_info : supports lists every variant with its category and modalities