Plot (mtb.plot)¶
Draw figures from a long results table. mtb.load_results returns one for
the stored tables. mtb.to_long and BatchResult.long return one for your
own runs.
| Name | Summary |
|---|---|
mtb.plot.bubble |
Draw the paper-style bubble table: methods as rows, metrics as columns. |
mtb.plot.bar |
Draw each method's overall score across datasets as a horizontal bar. |
mtb.plot.build_table |
Compute the ranks, scores and row order behind a bubble figure. |
mtb.plot.BubbleTable |
The numbers behind a bubble figure, as mtb.plot.build_table returns. |
mtb.plot.bubble
¶
bubble(
long_df,
*,
metrics=None,
methods=None,
order=None,
aggregate="dataset",
cmap=None,
title=None,
save=None,
show_language=True,
require_complete=False,
overall="rank",
na="warn",
)
Draw the paper-style bubble table: methods as rows, metrics as columns.
Metrics are grouped into task families, each led by an Overall bar;
rows run best first unless order is given. mtb.plot.build_table
returns the same numbers without drawing.
| PARAMETERS | DESCRIPTION |
|---|---|
long_df
|
Long table with
type
|
metrics
|
Metric codes to show, in this order within each family;
type
|
methods
|
Methods (rows) to show;
type
|
order
|
Methods to put first, in this order. The rest follow, best first. To
drop methods, use
type
|
aggregate
|
type
|
cmap
|
Matplotlib colormap for the first family;
type
|
title
|
Figure title;
type
|
save
|
File to write the figure to (tight bounding box); the suffix picks
the format, e.g.
type
|
show_language
|
Draw the
type
|
require_complete
|
With
type
|
overall
|
Across-dataset Overall:
type
|
na
|
How to report
type
|
| RETURNS | DESCRIPTION |
|---|---|
Figure
|
The figure (one axes), already saved when |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
|
ValueError
|
Duplicate |
ValueError
|
No method is complete under |
| WARNS | DESCRIPTION |
|---|---|
UserWarning
|
|
UserWarning
|
|
UserWarning
|
|
UserWarning
|
A dataset holds one method, or no method spans two datasets. |
UserWarning
|
A column has the same value in every row, or there is one method. |
UserWarning
|
Rows scored with the igraph Leiden backend are shown with stored rows of their dataset. |
Examples
>>> import multibench as mtb
>>> df = mtb.load_results("vertical", dataset="D11")
>>> fig = mtb.plot.bubble(df, metrics=["ARI", "NMI", "ASW"], title="D11",
... save="d11.pdf")
>>> multi = mtb.load_results("diagonal", dataset=["D24", "D25", "D28"])
>>> fig = mtb.plot.bubble(multi, aggregate="summary", require_complete=True)
Notes
bubble draws the table that mtb.plot.build_table computes. The
Notes of build_table give the Overall formulas, missing cells, row
order, input columns, name matching, errors and warnings.
Reading the figure. What each mark encodes:
- Metric circle (
aggregate="dataset"): radius = within-column rank,0.85 * sqrt(rank / n)withn= the methods scored in that column (n/a cells excluded), so the largest is the best; fill = the min-max scaled value on the family's colour ramp. The lightest fill is the lowest value of that column in this figure, not zero. - Metric bar (
aggregate="summary"): length and fill = the min-max scaled mean rank. - Score legend: the colour ramp, labelled "(scaled per column)":
LowandHighare the lowest and highest value of each column. - Family Overall bar: length = the family score min-max scaled across
the rows; fill = the score itself (the two differ only under
overall="mean_overall"). - Rows: best mean Overall first;
ordermoves rows only. - Rank legend: 1 = best, whereas
BubbleTable.ranks/FamilyBlock.ranksstore max-ranks (n= best). - Grey fill: every row of that column holds the same value, or the figure has one method, so there is nothing to compare.
- Footnote: the Overall formula in use, the grey columns, the
n/arule when a dash is drawn, and the chip key.
Chips and badge (show_language).
- Chip:
Py/R= language;?= a name the package does not know, such as your own or a renamed re-run. Lbadge: the method uses cell-type labels, and its clustering scores are not comparable with unsupervised rows. Aneeds_labelscolumn decides first. Otherwise the badge follows the frame'scategory: scMoMaT has it in mosaic only. With several categories, it shows when the method uses labels in any of them.
See Also
mtb.plot.build_table : the numbers behind the figure, to audit first.
mtb.plot.bar : one bar per method across datasets; same overall= formulas.
mtb.to_long : reshape mtb.evaluate's scores into a long table.
mtb.load_results : stored metric tables as a long table.
mtb.plot.bar
¶
bar(
long_df: DataFrame,
*,
metrics=None,
group: str | None = None,
top: int | None = None,
title: str | None = None,
cmap: str = "Blues",
show_datasets: bool = True,
save: str | None = None,
overall: str = "mean_overall",
)
Draw each method's overall score across datasets as a horizontal bar.
| PARAMETERS | DESCRIPTION |
|---|---|
long_df
|
Long table with
type
|
metrics
|
Metric codes to score;
type
|
group
|
Score one metric family only (the benchmark's two summary panels);
type
|
top
|
Keep the
type
|
title
|
Figure title;
type
|
cmap
|
Matplotlib colormap for the bars;
type
|
show_datasets
|
Overlay one dot per dataset's score on each bar; drawn only under
type
|
save
|
File to write the figure to (140 dpi, tight bounding box).
type
|
overall
|
Across-dataset Overall:
type
|
| RETURNS | DESCRIPTION |
|---|---|
Figure
|
One horizontal bar per method, best on top. |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
|
ValueError
|
Unknown |
ValueError
|
|
| WARNS | DESCRIPTION |
|---|---|
UserWarning
|
Several datasets and a method missing from some of them. |
UserWarning
|
A dataset holds one method, or no method spans two datasets. |
UserWarning
|
Rows scored with the igraph Leiden backend are shown with stored rows of their dataset. |
Examples
>>> import multibench as mtb
>>> # every dataset of the category that has a stored table
>>> df = mtb.load_results("diagonal")
>>> fig = mtb.plot.bar(df, group="clustering", top=10, save="clustering.png")
>>> fig = mtb.plot.bar(df, group="batch")
>>> fig = mtb.plot.bar(df, overall="rank") # bubble's default formula
Notes
Whiskers, dots and x label.
- Whiskers (
overall="mean_overall"): the SD of the per-dataset scores; a method present in one dataset gets none - there is no spread to show. - Dots (
overall="mean_overall",show_datasets): one dot per dataset score. - Under
overall="rank"the bar is not a mean of per-dataset scores, so neither whiskers nor dots are drawn. - X label: the single dataset, or the formula and the number of datasets.
Reading the score. It is rank-based, so it is only meaningful
relative to the other methods in the same figure. A method lacking a
metric is compared on the metrics it has under "mean_overall";
under "rank" the missing cell is rank 0 in that dataset.
Input. Concatenate several datasets' frames to summarise across
them; mtb.to_long and the BatchResult.long property give the same
frame for your own runs. metrics is case- and alias-tolerant
("ari" -> "ARI").
Shared rules. The Overall formulas, the tie-break shared with
mtb.plot.bubble and the warnings are those of
mtb.plot.build_table with aggregate="summary"; its Notes give
them.
Errors. An unknown metrics code gets a did-you-mean hint and the
list of metrics present. Batch metrics need a multi-batch dataset: a
single-batch design has none to compute, which is what the
group="batch" error says. An empty frame's error points at
mtb.load_results, which may have returned nothing (see its
UserWarning).
See Also
mtb.plot.bubble : per-dataset (or rank-averaged) bubble table, same overall= formulas.
mtb.load_results : stored metric tables as a long frame.
mtb.plot.build_table
¶
build_table(
long_df: DataFrame,
*,
metrics=None,
methods=None,
order=None,
aggregate: str = "dataset",
require_complete: bool = False,
overall: str = "rank",
na: str = "warn",
) -> BubbleTable
Compute the ranks, scores and row order behind a bubble figure.
Returns the numbers mtb.plot.bubble draws, without drawing.
| PARAMETERS | DESCRIPTION |
|---|---|
long_df
|
Long table with
type
|
metrics
|
Metric codes to keep, in this order within each family;
type
|
methods
|
Methods (rows) to keep;
type
|
order
|
Methods to put first, in this order. The rest follow, best first. To
drop methods, use
type
|
aggregate
|
type
|
require_complete
|
With
type
|
overall
|
Formula for each family's Overall under
type
|
na
|
How to report
type
|
| RETURNS | DESCRIPTION |
|---|---|
BubbleTable
|
Read |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
|
ValueError
|
Duplicate |
ValueError
|
No method is complete under |
| WARNS | DESCRIPTION |
|---|---|
UserWarning
|
|
UserWarning
|
|
UserWarning
|
|
UserWarning
|
A dataset holds one method, or no method spans two datasets. |
UserWarning
|
A column has the same value in every row, or there is one method. |
UserWarning
|
Rows scored with the igraph Leiden backend are shown with stored rows of their dataset. |
Examples
>>> import multibench as mtb
>>> tbl = mtb.plot.build_table(mtb.load_results("vertical", dataset="D11"))
>>> tbl.methods # rows, best first
>>> tbl.ranks # max-ranks per metric (n = best)
>>> multi = mtb.load_results("diagonal", dataset=["D24", "D25", "D28"])
>>> tbl = mtb.plot.build_table(multi, aggregate="summary",
... require_complete=True)
>>> tbl.coverage # datasets per method
Notes
Missing cells. A metric not computed for a method is an n/a
cell, drawn as a dash.
aggregate="dataset": the family Overall averages the ranks of the metrics the method has, and a column's ranks count only the methods scored in it.aggregate="summary": the cell is rank 0 in that dataset (the paper's rule), in the metric columns and in theoverall="rank"Overall;overall="mean_overall"skips it.
na sets how this is reported. na_cells lists the cells, e.g.
"YukiNet: DR and clustering Overall over 3 of 4 metrics (cLISI n/a)".
Overall formulas. overall= sets the family Overall under
aggregate="summary"; under "dataset" it is always minmax(mean
over metrics of max-rank). The two can order methods differently on
the same frame.
"rank"(bubble's default):minmax(mean over metrics of max-rank(mean over datasets of within-dataset max-rank))- the per-dataset ranks are averaged per metric, re-ranked across methods, averaged over metrics and min-max scaled. A method absent from a dataset scores rank 0 there (the paper's summary rule), which pulls it down."mean_overall"(bar's default):mean over datasets of minmax(mean over metrics of within-dataset max-rank)- each dataset gets its own min-max-scaled overall, and these are averaged; a dataset the method lacks is skipped.
Row order. The combined Overall is the mean of the family Overalls,
sorted best first with a stable sort, so tied methods keep alphabetical
order. Ranks and scores always come from the whole filtered frame;
order only moves rows.
Bubble and bar. mtb.plot.bar uses the same formulas and
tie-break. With aggregate="summary", the same overall= and the
metrics of one family (e.g. against bar(group="clustering")), both
figures order methods identically. Across both families they can
differ: bubble averages the family Overalls, bar scores all metrics
together.
require_complete. One UserWarning names each dropped method and
the datasets it lacks ("require_complete=True dropped 1 method ...:
MyRandom lacks D52s."). It has no effect under
aggregate="dataset".
A new dataset. A figure compares methods only where they share a
dataset, and the stored tables hold only the demo datasets. A
UserWarning names a dataset that holds one method, and says so when
no method spans two of the datasets. Plot such a dataset on its own, or
score your method on the demo dataset of its category and add that row.
A method alone on its dataset gets an Overall of 1.0 there under
"mean_overall". Under "dataset", the several-datasets warning
suggests aggregate="summary" only when every method has rows in at
least two datasets and no dataset holds a single method.
No comparison. A column whose rows all hold the same value, and
every column of a one-method figure, compares nothing: one
UserWarning names the columns, and the figure draws them in grey.
Ranks and scores stay as computed.
Leiden backend. The stored tables were clustered with leidenalg; the
igraph default can move ARI by up to about 0.1. When rows whose
scored_with starts with igraph/ meet stored rows of the same
dataset, and ARI, NMI or iF1 is shown, one UserWarning names those
methods and the fix: set mtb.config.DEFAULT.leiden_flavor =
"leidenalg" before mtb.evaluate.
Input columns. Only method, metric and value are
required.
dataset- groups rows foraggregate="summary"(absent = one dataset) and is part of the duplicate-row key. A"dataset"figure that mixes datasets averages them per method and adds a dataset cue to each row label (Name · D11orName · 3 ds).category- a single value makes theLbadge follow that category's variants.needs_labels(bool) - overrides theLbadge per method; the only way to badge a method the package does not know. NaN = no override.- Rows whose
metricis NaN are dropped.
To draw your own runs next to the stored table, concatenate the frames:
pd.concat([mtb.load_results("vertical", dataset="D11"), res.long]),
with res from mtb.run_all.
Name matching. metrics, methods and order match the
frame exactly, by canonical form ("ari" -> "ARI") or
case-insensitively; the frame's own spelling is kept. The family blocks
always stay in paper order; metrics outside the two paper families form
a neutral purple "Other" block.
Errors. A frame that looks like mtb.evaluate's wide output gets a
hint to convert it with mtb.to_long first; an unknown name gets a
did-you-mean hint and the values present. Duplicate rows raise
ValueError. Remove them, or rename each version (as mtb.sweep
does). A method with both True and False
needs_labels rows, or an empty frame, also raises ValueError.
See Also
mtb.plot.bubble : builds the table and draws the figure in one call. mtb.plot.BubbleTable : what is returned, field by field.
mtb.plot.BubbleTable
dataclass
¶
BubbleTable(
methods: list,
blocks: list,
matrix: DataFrame,
raw: DataFrame,
overall: Series,
aggregate: str = "dataset",
overall_basis: str = "rank",
datasets: tuple = (),
coverage: object = None,
method_datasets: object = None,
category: object = None,
needs_labels: object = None,
na_cells: object = None,
)
The numbers behind a bubble figure, as mtb.plot.build_table returns.
Rows are methods; the per-family matrices live in blocks.
mtb.plot.bubble draws the same figure from the long table.
| ATTRIBUTES | DESCRIPTION |
|---|---|
methods |
Row order: best first, or the
type
|
blocks |
One block per metric family present, in paper order, each with its
own
type
|
matrix |
Per-column min-max values (method x metric) of all families, in
figure order; the same object as
type
|
raw |
Unscaled matrix (method x metric) in figure order: metric means, or
mean within-dataset max-ranks under
type
|
overall |
Combined Overall per method (mean of the family Overalls); it sets
the row order unless
type
|
aggregate |
type
|
overall_basis |
Formula behind the family Overall under
type
|
datasets |
Dataset ids in the frame, sorted;
type
|
coverage |
Number of datasets each method has rows in;
type
|
method_datasets |
type
|
category |
The frame's single
type
|
needs_labels |
type
|
na_cells |
The
type
|
Examples
>>> import multibench as mtb
>>> tbl = mtb.plot.build_table(mtb.load_results("vertical", dataset="D11"))
>>> tbl.methods[:3] # the best three methods
>>> tbl.ranks.loc[tbl.methods[0]] # the best method's max-rank per metric
>>> tbl.blocks[0].overall # first family's Overall per method
Notes
Ranks. ranks and FamilyBlock.ranks hold per-column
max-ranks: n = best, ties share the higher number (R's
ties.method="max"). A method with no value in a column is NaN there,
and the column's n counts only the scored methods. The figure's
Rank legend counts 1 = best instead.
Ties. Values that agree to 9 decimal places count as equal in
ranks, norm and the Overall scores. An iLISI of 2.2e-16
therefore ties with 0.0.
Scaled values. matrix and norm are per-column min-max values
in [0, 1]; a constant (or all-NaN) column is all ones, as in the R code.
The figure draws a column whose rows all hold one value in grey, not at
the High end of the Score legend, and names it in the footnote. Every
column of a one-method figure is grey.
Blocks. Paper order: DR and clustering (blues), batch correction (greens), then "Other" (purples) for any metric outside the two.
See Also
mtb.plot.build_table : builds a table from a long table. mtb.plot.bubble : builds the table and draws the figure in one call.
norm
property
¶
All families' per-column min-max values, in figure order.
An alias of matrix: rows are methods, columns the metrics in
the order the figure draws them.
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
Method x metric values scaled to [0, 1] per column. |
ranks
property
¶
All families' per-column max-ranks, in figure order.
Rows are methods, columns the metrics in the order the figure
draws them; the class Notes give the max-rank convention.
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
Method x metric max-ranks: |