Skip to content

Changes

What each release adds and changes, newest first. A script that uses anything listed under New in 0.3.2 needs multibench-sc>=0.3.2. multibench --version prints the version you have.

New in 0.3.3

  • Method environments download faster. The environments over 2 GiB now come from the GitHub release in parts, instead of from Zenodo.
  • mtb.env.install prints its progress while it downloads, one line per tenth. A dropped or stalled connection resumes where it stopped.
  • cLISI and iLISI now compute on Colab and other Linux hosts with an older glibc. The LISI helper of scib is rebuilt there on first use.
  • The method scripts are fetched without git's per-file progress lines.
  • res.plot() and mtb.plot.bubble show their figure in a notebook without %matplotlib inline.
  • The tutorials install the environments and run the methods, on the demo data and on data written from an AnnData.

New in 0.3.2

Commands and flags

  • multibench fetch D46 D11 downloads demo datasets. --outputs downloads their stored run-all outputs instead, --scripts the method scripts, and --ref picks a commit or tag of the scripts.
  • multibench config lists the paths in use and where each one comes from. Its scripts_commit row is the commit of the method scripts. With MULTIBENCH_SCRIPTS_REF set, a scripts_ref row says whether it matches. --get data_path prints one value.
  • multibench info METHOD prints what one method needs: environment, GPU use, labels, ATAC input, variants and observed runtimes with cell counts.
  • multibench scan --strict exits with 1 when no requested row is runnable. With --methods, it exits with 1 when any named method has no runnable row. The error counts the blocked rows by cause: missing files, missing environment, no GPU on this computer, wrong ATAC kind, unreadable peak names and scripts not at MULTIBENCH_SCRIPTS_REF. It also exits with 1 while the method scripts are not fetched.
  • multibench scan --assume-gpu and multibench run-all --dry-run --assume-gpu check a GPU-node job from a login node without a GPU.
  • multibench scan --allow-atac-mismatch and multibench run-all --allow-atac-mismatch count a method as runnable when its ATAC file holds the other representation or peak names the method cannot read. The caveat stays.
  • multibench run-all --dry-run runs without --out-dir. Its commands then show <out_dir>.
  • multibench evaluate --leiden-flavor and multibench run-all --leiden-flavor set the Leiden backend of the scoring.
  • multibench run-all --batch CSV gives one batch id per cell, as mtb.run_all(batch=...) does. With --dry-run, a file of the wrong length exits with 1.
  • multibench evaluate --batch-column NAME and multibench run-all --batch-column NAME read the batch from one column of a --batch CSV with several columns, such as a sample sheet. A first column of cell ids still aligns the rows to the cells.
  • multibench evaluate --labels A.csv --labels B.csv stacks several label files in the order given, one batch per file. With --dataset, --category and --method, an order that contradicts the method's cell order exits with 1.
  • multibench evaluate --dataset D --method M --category C without --labels reads the label files of mtb.labels_for. --data-path sets their folder.
  • multibench evaluate --name NAME sets the row name of the long table, such as SCALEX_rerun. With --labels, --method can be left out. A --method given with --name must be a package method. It sets the label order, and a needs_labels column gives the rows the L badge of plot bubble when the method uses labels.
  • multibench convert --overwrite replaces files already in the folder.
  • multibench convert --batch-index N writes one file as batch N of a mosaic or cross dataset.
  • multibench convert --atac-from atac.h5ad --category diagonal reads the ATAC cells from a second file.
  • multibench plot bubble --na warn|skip|raise sets what happens to metrics a method lacks.
  • multibench run --runner "srun --gres=gpu:1 {env_cmd}": {env_cmd} is the command inside the method environment. mtb.run(cmd_template=...) takes it too.

Python arguments and settings

  • mtb.io.export_dataset(..., overwrite=True) replaces files already in the folder.
  • mtb.io.export_dataset(..., batch_index=N) writes the whole object as batch N of a mosaic or cross dataset.
  • mtb.io.export_dataset(rna, folder, atac=atac, atac_kind="peak", category="diagonal") writes RNA and ATAC from different cells.
  • A selector can end with a feature filter, such as rna="X[feature_types=Gene Expression]". multibench convert takes it too.
  • mtb.scan(..., assume_gpu=True) and mtb.run_all(..., dry_run=True, assume_gpu=True) skip this computer's GPU test.
  • mtb.scan(..., allow_atac_mismatch=True) and mtb.run_all(..., allow_atac_mismatch=True) do what --allow-atac-mismatch does.
  • mtb.labels_for(..., check=True) raises for a vertical or diagonal category on a folder of per-batch files. The default, None, warns.
  • mtb.recommend(..., atac="peak") filters by ATAC representation, as mtb.find_methods does.
  • MULTIBENCH_DATA_PATH and MULTIBENCH_REPO_PATH set data_path and repo_path. MULTIBENCH_SCRIPTS_REF picks a commit or tag of the method scripts. When the scripts present are at another commit, scan marks every row not runnable, run_all raises ValueError before any method starts, and the dry runs print a note. run_all(dry_run=True) prints the note only with verbose=True, as a [run_all] line on stdout after its count line. The config page lists every variable.
  • data_path, repo_path and envs_dir of mtb.config.DEFAULT accept a string.
  • mtb.load_batch(out_dir, data_path=...) says where the dataset folders of a saved result are. A data_path without the dataset folder raises ValueError. Without it, load_batch and rescore look for the dataset folder under the recorded data_path, then under data_root.

New fields in results

  • RunResult.obs_names holds the cell barcodes of the output rows.
  • RunResult.scripts_commit, env_flavor and hostname record the method scripts, the environment build (cpu, gpu or single) and the computer. Each run_all record in BatchResult.records has them too. A record of a reused output (skip_existing=True) copies them from the earlier record.
  • BatchResult.summary and summary.csv end with a caveat column and a reason column. reason says why a SKIPPED method did not run. Each record in BatchResult.records and batch_result.json has a caveat key. A record saved before 0.3.2 has none, and summary shows NaN there.
  • Each run_all record has data_root, the absolute path of the data folder it used. batch_result.json holds it too. load_batch and rescore write the folder they find into data_path and data_root.
  • run_all(batch=...) and run-all --batch save the batch in out_dir as batch_<hash>.csv. The record key batch_file names the file. BatchResult.save writes it too, also after rescore(batch=...). BatchResult.rescore() reuses that batch, also with labels=. It drops the batch, with a warning, when the dataset folder is not found. For an older result it warns once when metrics= includes a batch metric: pass batch= again to keep the batch metrics.
  • mtb.method_info(m)["gpu"] is required, used when present, not used or unknown.
  • reference_batch in each entry of mtb.method_info(m)["supports"] names a fixed reference batch, such as batch 3 for StabMap in cross.
  • The datasets column of mtb.recommend names the datasets each score comes from.
  • mtb.evaluate(...).attrs holds leiden_flavor, clustering, multibench_version and scib_version.
  • The scored_with column of mtb.to_long records how the scores were computed, such as leidenalg/sweep/0.3.2.

Behaviour changes in 0.3.2

These changes can stop a 0.3.1 script or change its result.

Dataset folders

  • mtb.io.export_dataset and multibench convert refuse to replace existing files. FileExistsError lists them. Pass overwrite=True or --overwrite. A failed call writes nothing.
  • category="mosaic" writes ATAC as atac<i>.h5, where 0.3.1 wrote atac_peak<i>.h5. Folders with the old names still work.
  • category="diagonal" pairs no barcodes, so RNA and ATAC may come from different cells. The labels go to rna_cty.csv and atac_cty.csv, not cty.csv.
  • batch= with category="vertical" or "diagonal" raises ValueError. 0.3.1 wrote per-batch files that no vertical method reads.
  • A missing label (NaN, None or '') raises ValueError.
  • RNA, ADT or peak values that are not whole numbers give a UserWarning. The methods expect raw counts.
  • mtb.io.to_canonical(src, folder, modality="gas") into a folder with atac_peak.h5 writes the rows in that file's cell order.

File checks

  • A vertical folder of per-batch files (rna1.h5, rna2.h5, ...) fails the file check. 0.3.1 read rna1.h5 as rna.h5.
  • Seurat_v5 needs rna.h5 and atac_peak.h5 from the same cells. A folder whose two files hold different cells, D28 included, fails its file check, and run_all skips Seurat_v5 there.
  • mtb.labels_for(..., "Seurat_v5") lists atac_cty before rna_cty, Seurat_v5's cell order.
  • mtb.labels_for(dataset, category, method) returns only the label files of the batches the method reads. UINMF on D52 gets cty1 and cty2, and the caveat of scan and run_all says UINMF reads batches 1-2 of 3. Batch 3 is not used.
  • In a diagonal folder, atac_gas.h5 must hold the cells of atac_peak.h5 in the same order. Otherwise the methods that read atac_gas.h5 fail the file check. Barcodes that differ only in a trailing -<n> count as the same cell.
  • UnitedNet reads cty.csv, where 0.3.1 read rna_cty.csv. inputs_for returns the key cty. A folder with only rna_cty.csv still works, and labels_for(..., "vertical", "UnitedNet") returns that file under cty. The input key rna_cty of mtb.run and multibench run --input still works, with a DeprecationWarning.
  • modalities=["rna", "atac_peak"] selects the methods that read peaks, and "atac_gas" the methods that read gene activity. In 0.3.1, find_methods and recommend returned every RNA+ATAC method for either token.
  • In scan and run_all, "atac_gas" no longer selects moETM, scMM and iPOLNG, which read peaks, unless methods= names them. Pass "atac_peak" for them. "atac_peak" also lists MultiMAP and Seurat_v3, which read both ATAC files.
  • A row whose ATAC file holds the other representation, such as peaks for Matilda, is not runnable in scan, and run_all skips it, also when methods= names the method. allow_atac_mismatch=True or --allow-atac-mismatch runs it with the caveat, as 0.3.1 did.
  • GLUE and Seurat_v3 are not runnable in scan when the peak file holds names without chromosome, start and end, such as peak_1. allow_atac_mismatch=True runs them anyway. 0.3.1 checked only GLUE on D28.

Runs

  • run_all into an out_dir that holds a saved result of the same dataset and category merges the records. Another dataset or category raises ValueError before any method runs. BatchResult.save does the same.
  • mtb.run(..., dry_run=True) checks the inputs as a real run does. A MuData or .h5mu input raises ValueError. An input file that does not exist gets only a note, such as SCALEX reads data/LUNG/atac_gas.h5, which does not exist.
  • mtb.run and multibench run, dry run included, apply the cell checks of the file check to the files you pass. Seurat_v5 inputs from different cells, or an atac_gas file whose cells differ from atac_peak in set or order, raise ValueError before anything runs.
  • GLUE reads a copy of atac_peak.h5 whose peak names mtb.run rewrites to chr:start-end, as for Seurat_v3. The command column of scan names that copy, which mtb.run writes first. Start GLUE with mtb.run, run_all or multibench run. The printed command alone fails in a job script.
  • mtb.run and multibench run warn once when an ATAC input holds the other representation or peak names the method cannot read, and still run. The dry run prints the same text on a line that starts with #.
  • run_all prints the caveat of each row it runs in its log, in the dry run too.
  • The result lines of run_all and rescore name the ARI, such as ARI 0.629. A FAIL or TIMEOUT line ends with the last line of the error that names a cause.
  • A real run_all logs skipping <method>: <reason> for a method it does not start, and records it with status SKIPPED and the reason in error. This covers each named method with no runnable row and each blocked method whose files are present. failures lists it only when methods= named it. A SKIPPED record never replaces an earlier run of that method in out_dir. When no requested method can start, run_all raises ValueError instead, and multibench run-all exits with 1.
  • scan and run_all, dry run included, raise ValueError before anything runs when methods= names a method with no variant in the category or modalities. 0.3.1 dropped that method. multibench run-all exits with 1.
  • BatchResult.rescore keeps FAIL and TIMEOUT records as they are. 0.3.1 scored them again, and a failed method came back as CHAIN_OK, RUN_OK_EVAL_FAILED or RUN_OK_NO_EMBEDDING.
  • A record that rescore cannot score, RUN_OK_NO_LABEL_MATCH or RUN_OK_EVAL_FAILED, keeps no metrics, batch_source or n_batches from an earlier scoring. 0.3.1 kept the batch columns, and the metrics when the output file could not be read.
  • BatchResult.rescore(metrics=...) with no ARI, NMI or iF1 keeps the stored label order and runs no Leiden sweep. 0.3.1 ranked the label orders again.
  • BatchResult.rescore(verbose=True) prints a line before each label-order ranking.
  • BatchResult.rescore no longer warns about the batch it takes from the label files when metrics= has no batch metric, such as rescore(metrics=["ARI", "NMI"]) on a result with several label files.
  • run_all(batch=...) takes the ids in the cell order of labels_for(dataset) and puts them in each method's cell order. 0.3.1 used the vector as given for every method. A Series or one-column DataFrame indexed by barcode is matched by barcode, in BatchResult.rescore(batch=...) too.
  • BatchResult.rescore(labels=...) aligns a Series or one-column DataFrame indexed by barcode, where 0.3.1 matched it by position. label_order then names the file order chosen.
  • A batch or labels CSV whose first column holds cell ids is aligned by that column. 0.3.1 matched its rows by position. This applies to run_all, rescore and run-all --batch, and to evaluate when the output has cell ids, also for several label files and --column. A column with no cell id is matched by position, with a UserWarning when it holds text. Numbers, such as R's row numbers, are matched by position unless they are exactly the cell ids. A first column with no header whose numbers restart at 0, as pd.concat(...).to_csv(path) writes, is read as an index. The same numbers out of order raise ValueError.
  • For batch=, labels= and these CSVs, ids that are missing, repeated or not cells of the dataset raise ValueError. run_all raises it before any method runs.
  • The first real run fills an existing empty repo_path with the method scripts, where 0.3.1 refused. multibench fetch --scripts does the same. For a repo_path that holds other files but no method scripts, runs raise RuntimeError and scan marks every row not runnable.
  • multibench evaluate also computes the batch metrics when --batch or several --labels are given, as mtb.evaluate does. --task is deprecated. Use --metrics.

Scores and figures

  • mtb.evaluate clips ASW, iASW, cLISI, ASW_batch, GC and iLISI to 0-1. A value within 1e-12 of 0 or 1 is recorded as 0 or 1.
  • Ranks and fills treat values that agree to 9 decimals as ties.
  • label_order_confidence stays within 0-1, because a runner-up ARI below 0 counts as 0. A value above 1 in an older folder reads 1.0 in summary. The column is numeric, also when every row is blank.
  • mtb.to_long adds the scored_with column. Scores with no record of how they were computed, such as a CSV read back, get unknown and a UserWarning.
  • mtb.plot.bubble draws a column whose rows all hold one value in grey, and every column of a one-method figure. The footnote and a UserWarning name them. Ranks and Overall are unchanged.
  • The chip key of bubble reads ? = a name the package does not know.
  • bubble widens a family pill to fit its header, and a figure of one to three metrics to fit its row labels and key. The height is unchanged.
  • bubble, build_table and bar warn about rows from datasets that share no method, a dataset with only one method, and igraph-scored rows next to stored rows of their dataset when ARI, NMI or iF1 is shown. bar also warns about an incomplete method x dataset table, as bubble does.
  • mtb.recommend warns when modalities or atac names a modality that none of the ranked datasets measured. With long_df=, it also warns when it ranks igraph-scored rows against stored rows of their dataset on ARI, NMI or iF1.
  • mtb.catalog.datasets() no longer has the empty columns assay, tissue, n_cells, n_batches and source.

Command line

  • multibench plot --input exits with 1 when --dataset or --methods removes all your rows, and warns when it removes some.
  • --format json writes / unescaped.
  • scan and run-all --dry-run add the atac and caveat columns to the short table when they apply.
  • run-all --assume-gpu without --dry-run is a usage error, exit code 2.
  • The dry runs of multibench run and run-all print # Dry run. Nothing was executed. first. run --dry-run then prints its notes, # multibench run would execute: and the command.
  • run-all exits with 3 after saving when a method failed, or a method named in --methods was skipped. 0.3.1 exited with 0. A chain such as run-all ... && plot ... now stops after run-all. A line on stderr names those methods, such as # 1 failed (Seurat_WNN), 1 skipped (totalVI). Other skipped methods go on a # Not run: ... line and do not set 3.
  • The last line of the No method can run error repeats the data path, modalities and other scan options of the call.
Changed messages

Many messages, warnings and printed lines were reworded. Apart from the download error under Downloads, the exception and warning types are the same. A script that matches message text may need an update.

  • Messages printed by multibench name its commands and flags, not Python calls.
  • caveat texts are sentences that start with the method name, such as totalVI needs raw counts. rna.h5 holds non-integer values. The reason column of scan is sentences with their own subjects, such as Environment scmb_r runs only on Linux, not on this computer. or scBridge needs atac_gas.h5, and LUNG has no such file.
  • inputs_for(..., check=True) and the files_reason column of scan say what the method needs and what the folder holds, such as SCALEX (diagonal) needs atac_gas.h5 in <folder>. The folder holds atac_cty.csv, atac_peak.h5, rna.h5 and rna_cty.csv. For MultiMAP and Seurat_v3, the text ends <method> reads both atac_peak.h5 and atac_gas.h5.
  • An unknown method id gets Unknown method stabmap. Did you mean StabMap? mtb.list_methods() shows all methods. A category the method does not run on gets SCALEX does not run on vertical data. Its categories: diagonal. multibench evaluate adds the --name and --labels flags for your own rows.
  • run_all with no method to run raises No method can run on D11 (vertical). With methods=, the message starts None of the requested methods (Matilda, totalVI) can run on D11 (vertical). On macOS and Windows, the error lists only the rows that something else also blocks.
  • The reason for scripts not at MULTIBENCH_SCRIPTS_REF ends with the fix, and scan --strict prints it in full: The method scripts are at 0c68f87, not deadbeef (MULTIBENCH_SCRIPTS_REF). Unset MULTIBENCH_SCRIPTS_REF, or set MULTIBENCH_REPO_PATH to a new folder and run multibench fetch --scripts.
  • A vertical folder whose ATAC file sits under atac_peak.h5 or atac_gas.h5 gets a reason that names the fix, such as MIRA reads atac.h5 for vertical. Rename atac_peak.h5 to atac.h5, or write it with category="vertical". When the file holds the other representation, it reads Matilda needs gene-activity ATAC (atac.h5), and the folder has peaks (atac_peak.h5).
  • The warnings about ATAC feature names give the fix, such as Only 0% of the ATAC feature names look like peaks such as chr1:100-200. If they are peaks, rename them to chr:start-end. They suggest gene activity only when the RNA names are unknown, as in to_canonical, or most ATAC names are RNA gene names.
  • A Series matched by position gets a warning that says how to check its order, such as The labels Series is matched by position, because the embedding has no cell ids. Check that it follows the embedding rows, ... When rescore cannot find the dataset folder, the warning names the folders it looked in and gives the fix, Pass data_path= to mtb.load_batch, or run rescore from the folder where run_all ran. A RUN_OK_NO_LABEL_MATCH record gets the same fix in its note.
  • Count lines say methods when each method has one row, and rows otherwise, such as [scan] 13 of 14 methods have their input files. So does the list header of the run_all error. mtb.run_all(dry_run=True) starts with [run_all] Dry run:. The commands header of run-all --dry-run explains its tags, such as [env missing] marks a row whose environment is not installed.
  • Reworded: the messages of export_dataset, to_canonical, inputs_for, params_for and evaluate. Also the environment messages, file checks, plot and scan warnings, the # Merged with line and the scan, convert and env help. Errors for an unknown modality or metrics= token, a missing folder or table, or a plot --input filter changed too.
  • scan with modalities=["rna", "atac_peak"] no longer warns that scBridge is left out. scBridge reads gene activity.
  • The barcode error for RNA and ATAC with no shared cells points to category="diagonal".
  • labels='obs:cell_type' on a MuData whose column is only in mdata['rna'].obs gets use labels='rna:cell_type'. The .h5mu recipes of multibench layout mosaic and multibench convert --help use --labels rna:cell_type.
  • The raw-count hints follow the selector you passed, such as adt='obsm:protein_counts'.
  • A method that needs a GPU on a computer without one gets <method> needs an NVIDIA GPU, and this computer has none.
  • On macOS and Windows, the refusal of mtb.run starts with Methods run only on Linux and names no install command. That of mtb.env.install starts with Method environments run only on Linux. The summary line of scan says Method environments run only on Linux. and what works on this computer. On Linux, it names mtb.env.doctor() when an environment is missing.
  • A count mismatch in mtb.evaluate names the argument and both counts, such as labels has 60 entries for 90 cells in the embedding. multibench evaluate names the flag and the file.
  • A per-cell CSV with several columns and none chosen gets the columns after the cell ids and how to choose one, such as The --batch file obs_sheet.csv has several columns after the cell ids: celltype, sample. Choose one with --batch-column. A column name that is not in the file gets the same list.
  • Missing batch labels get metrics='all' needs batch labels for ASW_batch, GC and iLISI. Pass batch=<vector>, or metrics='clustering'. A batch that no chosen metric uses gets batch= changes nothing here, because metrics=['ARI', 'NMI'] has no batch metric. ...
  • load_results with a result_path that does not exist raises FileNotFoundError naming the path. load_results("mosaic") starts with mosaic has no published tables. and names source='rerun'. multibench plot --category mosaic names --source rerun.
  • The warning of bubble and build_table for several datasets reads This figure averages each method over 3 datasets, D24, D25 and D28. Its rows mix datasets. multibench plot --input with --category prints # Added your 7 rows (MyMethod) to the stored diagonal table, which has 137 rows (source published).
  • bubble on your rows and stored rows under one method name says Your rows and the stored table both have uniPort on D28. Give your rows another name, such as uniPort_rerun, with to_long(method=...). multibench plot names evaluate --name instead.
  • scan and run_all name what does not match, such as Matilda does not run on cross data. Its categories: vertical. or The folder data/MYCITE does not exist. data holds D11, D28 and D52. ... A label file of the wrong length gets cty.csv has 15 labels, but rna.h5 has 20 cells. Give each cell one label, in the order of the cells. ...
  • An unknown method or metric in bubble, build_table, bar and multibench plot gets Unknown method Matlida. Did you mean Matilda? The table has A and B.
  • The warning of export_dataset for gene-activity ATAC with category="mosaic" starts with Every mosaic method reads peak ATAC. and names the fix.
  • mtb.catalog.metrics() describes cLISI and iLISI as the median over cells, scaled to 0-1.
  • The note before the clustering sweep of evaluate names the cell count, the metrics and the Leiden backend, and how to skip the sweep.
  • mtb.io.read_canonical of a missing path names mtb.config.DEFAULT.data_path.
  • mtb.describe_layout(category) prints only that category. The diagonal layout names each method on one line.
  • A failed download of the method scripts names the offline route: copy a fetched scripts folder and set MULTIBENCH_REPO_PATH.
  • env plan, env install and the install warning call an environment without a CPU build a single build, the same archive for CPU and GPU hosts. env status, env doctor and env plan show no flavor= for it.

Downloads

  • mtb.data.fetch and fetch_outputs raise OSError, naming the URL, when a download fails. 0.3.1 raised the urllib error, now kept as __cause__. Code that catches URLError must catch OSError.

Removed in 0.3.1

These calls no longer select or report anything and were removed.

Call What happens now
find_methods(..., available=...) TypeError: find_methods() got an unexpected keyword argument 'available'
method_info(m)["availability"], the "availability" key of mtb.env.plan() rows KeyError. The keys are gone.

Deprecated in 0.3.0

Renamed 0.2.1 spellings still work in 0.3.x with a DeprecationWarning and will be removed in 0.4. Removed ones raise an error.

0.2.1 0.3.0
mtb.plan(dataset, category, ...) mtb.scan(dataset, category, ...). It returns the same frame, with the command column.
mtb.plan_commands(...) mtb.scan(...)
mtb.runtime_hint(m) mtb.method_info(m)["runtime"]
mtb.plot.plot_bubble(...) mtb.plot.bubble(...)
evaluate(..., task="clustering" / "batch" / "all") evaluate(..., metrics="clustering" / "batch" / "all")
evaluate(..., family=...) evaluate(..., metrics=<token>)
evaluate(..., only=[...]) evaluate(..., metrics=[...])
load_results(..., method=...) load_results(..., methods=...)
load_results(..., metric=[...]) load_results(..., metrics=[...])
load_results(..., task=... / family=...) load_results(..., metrics=<token>)
recommend(..., task=... / family=...) recommend(..., metrics=<token>)
multibench evaluate --only A,B multibench evaluate --metrics A,B (stderr: warning: --only is deprecated. Use --metrics.)

Passing metrics= together with a deprecated selector raises TypeError: evaluate() got metrics= together with the deprecated ['task']; pass metrics= only.

To find old spellings in your code, run it with python -W error::DeprecationWarning. The warnings then become errors.

These spellings were removed in 0.3.0:

removed error
mtb.command_preview(...) AttributeError. Use mtb.run(..., dry_run=True).
mtb.io.from_mudata(...) AttributeError. Use mtb.io.export_dataset(mdata, ..., rna="rna", atac="atac", atac_kind="peak", labels="rna:celltype").
mtb.io.write_labels(...) AttributeError. export_dataset writes the label files. By hand, use pd.Series(labels, name="x").to_csv(path, index=False).
list_methods(task=...) / (runnable=...) TypeError: list_methods() only takes category since 0.3.0; task=... are find_methods filters - use find_methods(category, task=...)
find_methods("vertical", "clustering") (positional filters) TypeError: find_methods() takes from 0 to 1 positional arguments but 2 were given
method_info(m, files_dir=...) TypeError: method_info() got an unexpected keyword argument 'files_dir'. The catalog columns are in mtb.catalog.methods().
cite(methods=[...]) TypeError: cite() got an unexpected keyword argument 'methods'. Use cite([...]) or one id per argument.
cite([ids], "bibtex") TypeError: cite(): method ids must be strings, got list; pass one list (cite(['Matilda', 'MOFA2'])) or one id per argument. fmt= is keyword-only.
inputs_for(dataset, method, category) TypeError: inputs_for argument order is (dataset, category, method) since 0.3.0; you passed (dataset, method, category)
labels_for(dataset, method, category) TypeError: labels_for argument order is (dataset, category, method) since 0.3.0; you passed (dataset, method, category)
labels_for(dataset, "<data_path>") TypeError: labels_for: pass data_path= by keyword; the 2nd positional argument is category since 0.3.0 (order: (dataset, category, method))
inputs_for(..., ["rna", "adt"]) (positional modalities) TypeError: inputs_for() takes 3 positional arguments but 4 were given
scan(dataset, category, data_path) (positional) TypeError: scan() takes from 1 to 2 positional arguments but 3 were given
run(method, category, "clustering", ...) (positional task) TypeError: run() takes 2 positional arguments but 3 positional arguments (and 2 keyword-only arguments) were given
run_all(dataset, category) without out_dir TypeError: run_all() needs out_dir= for a real run (dry_run=True returns the scan frame without one)
BatchResult.rescore(only=...) TypeError: BatchResult.rescore() got an unexpected keyword argument 'only'. Use metrics=.
evaluate(output, category, ...) (positional category) TypeError: evaluate() got multiple values for argument 'labels'. labels is the second positional argument, and the rest are keyword-only.
evaluate(..., slow_metrics=True) TypeError: evaluate() got slow_metrics=, removed in 0.3.0: pass metrics=[...] without cLISI/iLISI (kBET is computed only when it is named in that list)
evaluate(..., column="x") TypeError: evaluate() got column=, removed in 0.3.0: pass the Series/column itself as labels= / batch= / clustering=
evaluate(..., metric_set="scib") TypeError: evaluate() got metric_set=, removed in 0.3.0: only the scIB metric set exists - drop the argument
to_long(df, method, dataset, category) (positional) TypeError: to_long() takes 1 positional argument but 4 were given
to_long(..., needs_labels=True) TypeError: to_long() got an unexpected keyword argument 'needs_labels'. Add the column to the frame.
load_results(category, "clustering") (positional) TypeError: load_results() takes from 0 to 1 positional arguments but 2 were given
load_results(..., metric_set="scib") TypeError: load_results() got metric_set=, removed in 0.3.0: only the scIB metric set exists - drop the argument
available_datasets(..., metric_set=... / clustering=...) TypeError: available_datasets() got an unexpected keyword argument 'metric_set' (likewise 'clustering')
mtb.config.category_folder / metric_set_dir, mtb.plot.render / FamilyBlock, mtb.env.group_for and the other mtb.env helpers Still importable, but internal: out of __all__ and hidden from dir().