Skip to content

Plot (mtb.plot)

Draw figures from a long results table. mtb.load_results returns one for the stored tables. mtb.to_long and BatchResult.long return one for your own runs.

Name Summary
mtb.plot.bubble Draw the paper-style bubble table: methods as rows, metrics as columns.
mtb.plot.bar Draw each method's overall score across datasets as a horizontal bar.
mtb.plot.build_table Compute the ranks, scores and row order behind a bubble figure.
mtb.plot.BubbleTable The numbers behind a bubble figure, as mtb.plot.build_table returns.

mtb.plot.bubble

bubble(
    long_df,
    *,
    metrics=None,
    methods=None,
    order=None,
    aggregate="dataset",
    cmap=None,
    title=None,
    save=None,
    show_language=True,
    require_complete=False,
    overall="rank",
    na="warn",
)

Draw the paper-style bubble table: methods as rows, metrics as columns.

Metrics are grouped into task families, each led by an Overall bar; rows run best first unless order is given. mtb.plot.build_table returns the same numbers without drawing.

PARAMETERS DESCRIPTION
long_df

Long table with method, metric, value and optional dataset columns, as from mtb.load_results, mtb.to_long or the BatchResult.long property (more in Notes).

type DataFrame

metrics

Metric codes to show, in this order within each family; None = every metric in the frame.

type list of str | None default None

methods

Methods (rows) to show; None = every method in the frame.

type list of str | None default None

order

Methods to put first, in this order. The rest follow, best first. To drop methods, use methods.

type list of str | None default None

aggregate

"dataset": one dataset's values, drawn as circles. "summary": within-dataset ranks averaged across datasets, drawn as bars (the paper's panel c).

type ('dataset', 'summary') default "dataset"

cmap

Matplotlib colormap for the first family; None = blues, greens and purples per family.

type str | None default None

title

Figure title; None = no title.

type str | None default None

save

File to write the figure to (tight bounding box); the suffix picks the format, e.g. .pdf, .png, .svg.

type str or path - like | None default None

show_language

Draw the Py / R chip and the L (supervised) badge left of each row, plus a key line (Notes).

type bool default True

require_complete

With aggregate="summary": keep only the methods present in every dataset.

type bool default False

overall

Across-dataset Overall: "rank" re-ranks mean ranks (missing dataset = rank 0); "mean_overall" averages per-dataset Overalls (missing dataset skipped). Applies under aggregate="summary".

type ('rank', 'mean_overall') default "rank"

na

How to report n/a cells (a method lacking a metric): warn once, stay silent, or raise.

type ('warn', 'skip', 'raise') default "warn"

RETURNS DESCRIPTION
Figure

The figure (one axes), already saved when save is given.

RAISES DESCRIPTION
ValueError

long_df lacks method / metric / value, or metrics / methods / order names an absent value.

ValueError

Duplicate (method[, dataset], metric) rows, or an invalid aggregate / overall / na.

ValueError

No method is complete under require_complete=True, or n/a cells under na="raise".

WARNS DESCRIPTION
UserWarning

n/a cells under na="warn", or several datasets under aggregate="dataset".

UserWarning

aggregate="summary" on an incomplete method x dataset matrix.

UserWarning

require_complete=True dropped methods; each is named with the datasets it lacks.

UserWarning

A dataset holds one method, or no method spans two datasets.

UserWarning

A column has the same value in every row, or there is one method.

UserWarning

Rows scored with the igraph Leiden backend are shown with stored rows of their dataset.

Examples

>>> import multibench as mtb
>>> df = mtb.load_results("vertical", dataset="D11")
>>> fig = mtb.plot.bubble(df, metrics=["ARI", "NMI", "ASW"], title="D11",
...                       save="d11.pdf")
>>> multi = mtb.load_results("diagonal", dataset=["D24", "D25", "D28"])
>>> fig = mtb.plot.bubble(multi, aggregate="summary", require_complete=True)
Notes

bubble draws the table that mtb.plot.build_table computes. The Notes of build_table give the Overall formulas, missing cells, row order, input columns, name matching, errors and warnings.

Reading the figure. What each mark encodes:

  • Metric circle (aggregate="dataset"): radius = within-column rank, 0.85 * sqrt(rank / n) with n = the methods scored in that column (n/a cells excluded), so the largest is the best; fill = the min-max scaled value on the family's colour ramp. The lightest fill is the lowest value of that column in this figure, not zero.
  • Metric bar (aggregate="summary"): length and fill = the min-max scaled mean rank.
  • Score legend: the colour ramp, labelled "(scaled per column)": Low and High are the lowest and highest value of each column.
  • Family Overall bar: length = the family score min-max scaled across the rows; fill = the score itself (the two differ only under overall="mean_overall").
  • Rows: best mean Overall first; order moves rows only.
  • Rank legend: 1 = best, whereas BubbleTable.ranks / FamilyBlock.ranks store max-ranks (n = best).
  • Grey fill: every row of that column holds the same value, or the figure has one method, so there is nothing to compare.
  • Footnote: the Overall formula in use, the grey columns, the n/a rule when a dash is drawn, and the chip key.

Chips and badge (show_language).

  • Chip: Py / R = language; ? = a name the package does not know, such as your own or a renamed re-run.
  • L badge: the method uses cell-type labels, and its clustering scores are not comparable with unsupervised rows. A needs_labels column decides first. Otherwise the badge follows the frame's category: scMoMaT has it in mosaic only. With several categories, it shows when the method uses labels in any of them.
See Also

mtb.plot.build_table : the numbers behind the figure, to audit first. mtb.plot.bar : one bar per method across datasets; same overall= formulas. mtb.to_long : reshape mtb.evaluate's scores into a long table. mtb.load_results : stored metric tables as a long table.

mtb.plot.bar

bar(
    long_df: DataFrame,
    *,
    metrics=None,
    group: str | None = None,
    top: int | None = None,
    title: str | None = None,
    cmap: str = "Blues",
    show_datasets: bool = True,
    save: str | None = None,
    overall: str = "mean_overall",
)

Draw each method's overall score across datasets as a horizontal bar.

PARAMETERS DESCRIPTION
long_df

Long table with method, metric, value columns and an optional dataset (absent = one dataset), e.g. from mtb.load_results.

type DataFrame

metrics

Metric codes to score; None = every metric. Ignored when group is given.

type list of str | None default None

group

Score one metric family only (the benchmark's two summary panels); None = the metrics selection.

type (None, 'clustering', 'batch') default None

top

Keep the top best methods, scored against all of them; None = every method.

type int | None default None

title

Figure title; None = the number of datasets. With group, (<group> metrics) is appended.

type str | None default None

cmap

Matplotlib colormap for the bars; group="batch" always uses "Greens" (the paper's family colour).

type str default 'Blues'

show_datasets

Overlay one dot per dataset's score on each bar; drawn only under overall="mean_overall" with several datasets.

type bool default True

save

File to write the figure to (140 dpi, tight bounding box).

type str | None default None

overall

Across-dataset Overall: "rank" re-ranks mean ranks (missing dataset = rank 0); "mean_overall" averages per-dataset Overalls (missing dataset skipped).

type ('mean_overall', 'rank') default "mean_overall"

RETURNS DESCRIPTION
Figure

One horizontal bar per method, best on top.

RAISES DESCRIPTION
ValueError

long_df is empty or lacks method / metric / value.

ValueError

Unknown metrics code, or an invalid group / overall.

ValueError

group names a family with no metric in the frame.

WARNS DESCRIPTION
UserWarning

Several datasets and a method missing from some of them.

UserWarning

A dataset holds one method, or no method spans two datasets.

UserWarning

Rows scored with the igraph Leiden backend are shown with stored rows of their dataset.

Examples

>>> import multibench as mtb
>>> # every dataset of the category that has a stored table
>>> df = mtb.load_results("diagonal")
>>> fig = mtb.plot.bar(df, group="clustering", top=10, save="clustering.png")
>>> fig = mtb.plot.bar(df, group="batch")
>>> fig = mtb.plot.bar(df, overall="rank")       # bubble's default formula
Notes

Whiskers, dots and x label.

  • Whiskers (overall="mean_overall"): the SD of the per-dataset scores; a method present in one dataset gets none - there is no spread to show.
  • Dots (overall="mean_overall", show_datasets): one dot per dataset score.
  • Under overall="rank" the bar is not a mean of per-dataset scores, so neither whiskers nor dots are drawn.
  • X label: the single dataset, or the formula and the number of datasets.

Reading the score. It is rank-based, so it is only meaningful relative to the other methods in the same figure. A method lacking a metric is compared on the metrics it has under "mean_overall"; under "rank" the missing cell is rank 0 in that dataset.

Input. Concatenate several datasets' frames to summarise across them; mtb.to_long and the BatchResult.long property give the same frame for your own runs. metrics is case- and alias-tolerant ("ari" -> "ARI").

Shared rules. The Overall formulas, the tie-break shared with mtb.plot.bubble and the warnings are those of mtb.plot.build_table with aggregate="summary"; its Notes give them.

Errors. An unknown metrics code gets a did-you-mean hint and the list of metrics present. Batch metrics need a multi-batch dataset: a single-batch design has none to compute, which is what the group="batch" error says. An empty frame's error points at mtb.load_results, which may have returned nothing (see its UserWarning).

See Also

mtb.plot.bubble : per-dataset (or rank-averaged) bubble table, same overall= formulas. mtb.load_results : stored metric tables as a long frame.

mtb.plot.build_table

build_table(
    long_df: DataFrame,
    *,
    metrics=None,
    methods=None,
    order=None,
    aggregate: str = "dataset",
    require_complete: bool = False,
    overall: str = "rank",
    na: str = "warn",
) -> BubbleTable

Compute the ranks, scores and row order behind a bubble figure.

Returns the numbers mtb.plot.bubble draws, without drawing.

PARAMETERS DESCRIPTION
long_df

Long table with method, metric, value columns; optional dataset, category and needs_labels columns (Notes).

type DataFrame

metrics

Metric codes to keep, in this order within each family; None = every metric in the frame.

type list of str | None default None

methods

Methods (rows) to keep; None = every method in the frame.

type list of str | None default None

order

Methods to put first, in this order. The rest follow, best first. To drop methods, use methods.

type list of str | None default None

aggregate

"dataset": raw metric values of one dataset (several are averaged per method). "summary": within-dataset max-ranks averaged across datasets (the paper's panel c).

type ('dataset', 'summary') default "dataset"

require_complete

With aggregate="summary": keep only the methods present in every dataset.

type bool default False

overall

Formula for each family's Overall under aggregate="summary": "rank" (the paper's panel rule) or "mean_overall" (bar's default); formulas in Notes.

type ('rank', 'mean_overall') default "rank"

na

How to report n/a cells (a method lacking a metric): "warn", "skip" (silent, nothing is dropped) or "raise"; also stored as na_cells.

type ('warn', 'skip', 'raise') default "warn"

RETURNS DESCRIPTION
BubbleTable

Read methods (row order: best first unless order is given), ranks (max-ranks, n = best) and overall; all fields are under mtb.plot.BubbleTable.

RAISES DESCRIPTION
ValueError

long_df lacks method / metric / value, or metrics / methods / order names an absent value.

ValueError

Duplicate (method[, dataset], metric) rows, or an invalid aggregate / overall / na.

ValueError

No method is complete under require_complete=True, or n/a cells under na="raise".

WARNS DESCRIPTION
UserWarning

n/a cells under na="warn", or several datasets under aggregate="dataset".

UserWarning

aggregate="summary" on an incomplete method x dataset matrix.

UserWarning

require_complete=True dropped methods; each is named with the datasets it lacks.

UserWarning

A dataset holds one method, or no method spans two datasets.

UserWarning

A column has the same value in every row, or there is one method.

UserWarning

Rows scored with the igraph Leiden backend are shown with stored rows of their dataset.

Examples

>>> import multibench as mtb
>>> tbl = mtb.plot.build_table(mtb.load_results("vertical", dataset="D11"))
>>> tbl.methods                      # rows, best first
>>> tbl.ranks                        # max-ranks per metric (n = best)
>>> multi = mtb.load_results("diagonal", dataset=["D24", "D25", "D28"])
>>> tbl = mtb.plot.build_table(multi, aggregate="summary",
...                            require_complete=True)
>>> tbl.coverage                     # datasets per method
Notes

Missing cells. A metric not computed for a method is an n/a cell, drawn as a dash.

  • aggregate="dataset": the family Overall averages the ranks of the metrics the method has, and a column's ranks count only the methods scored in it.
  • aggregate="summary": the cell is rank 0 in that dataset (the paper's rule), in the metric columns and in the overall="rank" Overall; overall="mean_overall" skips it.

na sets how this is reported. na_cells lists the cells, e.g. "YukiNet: DR and clustering Overall over 3 of 4 metrics (cLISI n/a)".

Overall formulas. overall= sets the family Overall under aggregate="summary"; under "dataset" it is always minmax(mean over metrics of max-rank). The two can order methods differently on the same frame.

  • "rank" (bubble's default): minmax(mean over metrics of max-rank(mean over datasets of within-dataset max-rank)) - the per-dataset ranks are averaged per metric, re-ranked across methods, averaged over metrics and min-max scaled. A method absent from a dataset scores rank 0 there (the paper's summary rule), which pulls it down.
  • "mean_overall" (bar's default): mean over datasets of minmax(mean over metrics of within-dataset max-rank) - each dataset gets its own min-max-scaled overall, and these are averaged; a dataset the method lacks is skipped.

Row order. The combined Overall is the mean of the family Overalls, sorted best first with a stable sort, so tied methods keep alphabetical order. Ranks and scores always come from the whole filtered frame; order only moves rows.

Bubble and bar. mtb.plot.bar uses the same formulas and tie-break. With aggregate="summary", the same overall= and the metrics of one family (e.g. against bar(group="clustering")), both figures order methods identically. Across both families they can differ: bubble averages the family Overalls, bar scores all metrics together.

require_complete. One UserWarning names each dropped method and the datasets it lacks ("require_complete=True dropped 1 method ...: MyRandom lacks D52s."). It has no effect under aggregate="dataset".

A new dataset. A figure compares methods only where they share a dataset, and the stored tables hold only the demo datasets. A UserWarning names a dataset that holds one method, and says so when no method spans two of the datasets. Plot such a dataset on its own, or score your method on the demo dataset of its category and add that row.

A method alone on its dataset gets an Overall of 1.0 there under "mean_overall". Under "dataset", the several-datasets warning suggests aggregate="summary" only when every method has rows in at least two datasets and no dataset holds a single method.

No comparison. A column whose rows all hold the same value, and every column of a one-method figure, compares nothing: one UserWarning names the columns, and the figure draws them in grey. Ranks and scores stay as computed.

Leiden backend. The stored tables were clustered with leidenalg; the igraph default can move ARI by up to about 0.1. When rows whose scored_with starts with igraph/ meet stored rows of the same dataset, and ARI, NMI or iF1 is shown, one UserWarning names those methods and the fix: set mtb.config.DEFAULT.leiden_flavor = "leidenalg" before mtb.evaluate.

Input columns. Only method, metric and value are required.

  • dataset - groups rows for aggregate="summary" (absent = one dataset) and is part of the duplicate-row key. A "dataset" figure that mixes datasets averages them per method and adds a dataset cue to each row label (Name · D11 or Name · 3 ds).
  • category - a single value makes the L badge follow that category's variants.
  • needs_labels (bool) - overrides the L badge per method; the only way to badge a method the package does not know. NaN = no override.
  • Rows whose metric is NaN are dropped.

To draw your own runs next to the stored table, concatenate the frames: pd.concat([mtb.load_results("vertical", dataset="D11"), res.long]), with res from mtb.run_all.

Name matching. metrics, methods and order match the frame exactly, by canonical form ("ari" -> "ARI") or case-insensitively; the frame's own spelling is kept. The family blocks always stay in paper order; metrics outside the two paper families form a neutral purple "Other" block.

Errors. A frame that looks like mtb.evaluate's wide output gets a hint to convert it with mtb.to_long first; an unknown name gets a did-you-mean hint and the values present. Duplicate rows raise ValueError. Remove them, or rename each version (as mtb.sweep does). A method with both True and False needs_labels rows, or an empty frame, also raises ValueError.

See Also

mtb.plot.bubble : builds the table and draws the figure in one call. mtb.plot.BubbleTable : what is returned, field by field.

mtb.plot.BubbleTable dataclass

BubbleTable(
    methods: list,
    blocks: list,
    matrix: DataFrame,
    raw: DataFrame,
    overall: Series,
    aggregate: str = "dataset",
    overall_basis: str = "rank",
    datasets: tuple = (),
    coverage: object = None,
    method_datasets: object = None,
    category: object = None,
    needs_labels: object = None,
    na_cells: object = None,
)

The numbers behind a bubble figure, as mtb.plot.build_table returns.

Rows are methods; the per-family matrices live in blocks. mtb.plot.bubble draws the same figure from the long table.

ATTRIBUTES DESCRIPTION
methods

Row order: best first, or the order methods first; methods.index(name) + 1 is the row's position (1 = top).

type list of str

blocks

One block per metric family present, in paper order, each with its own raw, norm, ranks and overall.

type list of FamilyBlock

matrix

Per-column min-max values (method x metric) of all families, in figure order; the same object as norm.

type DataFrame

raw

Unscaled matrix (method x metric) in figure order: metric means, or mean within-dataset max-ranks under "summary".

type DataFrame

overall

Combined Overall per method (mean of the family Overalls); it sets the row order unless order is given.

type Series

aggregate

"dataset" (metric markers are circles) or "summary" (bars, the paper's panel c).

type str

overall_basis

Formula behind the family Overall under "summary": "rank" or "mean_overall" (the overall= of mtb.plot.bubble).

type str

datasets

Dataset ids in the frame, sorted; () without a dataset column.

type tuple of str

coverage

Number of datasets each method has rows in; "summary" only, else None.

type Series or None

method_datasets

{method: [dataset, ...]} from the frame; None without a dataset column.

type dict or None

category

The frame's single category value, else None (mixed or absent); drives the supervised L badge.

type str or None

needs_labels

{method: bool} from an optional needs_labels column, empty without it; overrides the package's own flag for the L badge.

type dict or None

na_cells

The n/a cells: one line per family and method (per dataset and method under "summary").

type list of str or None

Examples

>>> import multibench as mtb
>>> tbl = mtb.plot.build_table(mtb.load_results("vertical", dataset="D11"))
>>> tbl.methods[:3]                  # the best three methods
>>> tbl.ranks.loc[tbl.methods[0]]    # the best method's max-rank per metric
>>> tbl.blocks[0].overall            # first family's Overall per method
Notes

Ranks. ranks and FamilyBlock.ranks hold per-column max-ranks: n = best, ties share the higher number (R's ties.method="max"). A method with no value in a column is NaN there, and the column's n counts only the scored methods. The figure's Rank legend counts 1 = best instead.

Ties. Values that agree to 9 decimal places count as equal in ranks, norm and the Overall scores. An iLISI of 2.2e-16 therefore ties with 0.0.

Scaled values. matrix and norm are per-column min-max values in [0, 1]; a constant (or all-NaN) column is all ones, as in the R code. The figure draws a column whose rows all hold one value in grey, not at the High end of the Score legend, and names it in the footnote. Every column of a one-method figure is grey.

Blocks. Paper order: DR and clustering (blues), batch correction (greens), then "Other" (purples) for any metric outside the two.

See Also

mtb.plot.build_table : builds a table from a long table. mtb.plot.bubble : builds the table and draws the figure in one call.

norm property

norm: DataFrame

All families' per-column min-max values, in figure order.

An alias of matrix: rows are methods, columns the metrics in the order the figure draws them.

RETURNS DESCRIPTION
DataFrame

Method x metric values scaled to [0, 1] per column.

ranks property

ranks: DataFrame

All families' per-column max-ranks, in figure order.

Rows are methods, columns the metrics in the order the figure draws them; the class Notes give the max-rank convention.

RETURNS DESCRIPTION
DataFrame

Method x metric max-ranks: n = best, NaN where the method has no value for that metric.