Skip to content

PR8/8: Add comparison reports in hwgraph#127

Open
ErwanAliasr1 wants to merge 27 commits into
mainfrom
hwgraph-comparison-reports
Open

PR8/8: Add comparison reports in hwgraph#127
ErwanAliasr1 wants to merge 27 commits into
mainfrom
hwgraph-comparison-reports

Conversation

@ErwanAliasr1

Copy link
Copy Markdown
Collaborator

This PR adds :

  • add a traces-comparison report and exec-summary scorecard
  • add a per-scaling-step traces comparison under scaling/summary/

ErwanAliasr1 and others added 2 commits July 15, 2026 16:46
The turbostat "starting/stopping background monitoring" messages were printed
for every single job, making the output very verbose and hard to read on runs
with many jobs. This information is still useful for debugging, so rather than
removing it, it is now gated behind a new --verbose (-v) flag (off by default).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
Pass --yaml to every stress-ng stressor so it writes its metrics next to
the stdout file (as <name>.yaml). We don't parse it yet, but capturing the
structured YAML output gives us a machine-readable source to build richer
parsing features on later, rather than scraping the human-oriented stdout.

Add a parsing fixture for stress-ng 0.21.03 under
hwbench/tests/parsing/stressng/v02103b/, including the real .yaml artifact
alongside stdout, so the auto-discovered parsing tests now cover this
version.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
@ErwanAliasr1
ErwanAliasr1 force-pushed the hwgraph-comparison-reports branch from 69a3a60 to bdccfbd Compare July 15, 2026 16:48
ErwanAliasr1 and others added 25 commits July 16, 2026 12:22
Benchmark output files -- stdout, stderr, version-stdout, version-stderr
and, for stress-ng, the --yaml metrics file -- were named after the engine's
`name` (e.g. "stressngint128"). That name is shared by every run a single
job expands into: the scaling matrix (hosting_cpu_cores_scaling /
stressor_range) turns one job into many runs that all reuse it. As they
share one output directory, each iteration overwrote the previous one's
files, leaving only the last iteration on disk.

Introduce ExternalBench.output_basename, which prefixes the engine `name`
with get_name_with_position() -- the per-benchmark id already used as the
results.json key (and the expanded-job section header). _write_output and
the stress-ng --yaml file now use it, so files look like
"cpu10_5_stressngint128-stdout" / "....yaml" and map one-to-one to their
results.json entry. `name` keeps its original meaning; External falls back
to it for non-benchmark commands (External_Simple).

While here, fix a latent bug this path exposes: get_name_with_position()
guards on `if not self.benchmark`, but BenchmarkParameters never initialised
that attribute (only set_benchmark() did), so reading it before the owning
Benchmark is attached raised AttributeError instead of falling back to the
bare job name. Initialise self.benchmark = None.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
Add a per-NUMA-domain view of the per-core CPU metrics (frequency, IPC and
core power): instead of hundreds of unreadable per-core lines, plot one
averaged line per NUMA domain. Following the per-core convention, each metric
is rendered as "all_numa" (every core of each domain) and "pinned_numa" (only
the cores pinned during the job, grouped by their domain, dropping domains
with no pinned core; annotated with the pinned range).

This requires the NUMA topology in the trace, which hwbench collected but did
not serialize. cpu.dump() now writes:
  - numa_nodes: {domain: [logical cores]}
  - numa_distances: the inter-node distance matrix (kept for future
    topology-aware features)
Both are parsed from numactl -H; NUMA.get_distances() exposes the matrix that
was previously only used to derive quadrants. Traces predating this change
lack the topology and simply skip the NUMA graphs.

Tests reuse the existing NUMA fixtures: distance-matrix assertions in the
parse tests, a cpu.dump() topology test on the mocked hardware.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
Complement the per-NUMA-domain line graphs with two heatmaps:

- A per-host NUMA distance heatmap (domain x domain, color = inter-domain
  distance), rendered once per trace in environment/by_host/<host>/, so the
  topology (which domains are close/far) is visible independently of any
  benchmark.

- A per-metric NUMA domain x time heatmap for frequency and IPC, next to the
  corresponding line graph in all_numa/pinned_numa: Y = NUMA domains, X = time
  (as in the line graphs), color = the domain's value at each step. A legend
  box on the left lists each domain's cores condensed via cpu_list_to_range()
  (individually bracketed, right-aligned, centered dash), like the component
  legend of the other graphs.

Graph.render() gains an extra_legend parameter so a manually placed legend is
included in the tight bounding box.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
IPC is usually interesting to see how cores are performing.
The frequency is not always enough to get a good estimation of the
core's performance, so let's report the IPC too.

Add an "ipc" SMP scaling graph, mirroring the cpu_clock one: add_perf() gains
ipc/ipc_err/ipc_cores parameters that aggregate the min/mean/max IPC across
cores (optionally restricted to the pinned cores), and smp_scaling renders an
ipc graph type with the same all_cores/pinned_cores variants and errorbars.

IPC is not always collected, so it is only aggregated and rendered when the
trace actually has an IPC context; IPC-less traces are unaffected.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
Compare traces per NUMA domain across the scaling sweep. The first trace is
the reference; one heatmap is produced for each other trace, with the compared
trace name in the filename (scaling_<type>_numa_delta_<ref>_vs_<other>_<engine>).

Y = NUMA domains, X = scaling step (worker count), color = signed delta
value(reference) - value(other): red when the compared trace is lower than the
reference, green when higher, white when equal (diverging map centered at 0).
Rendered for frequency (cpu_clock) and IPC (cpu_ipc, when present), each in the
all_cores and (when the sweep pins cores) pinned_cores directories next to the
line graphs. On the pinned view, domains not pinned at a given step have no
data and are drawn black.

This will help understanding what NUMA mode was more or less active
during a step of the scaling.

A left box lists each domain's cores (condensed, via numa_core_blocks) and the
colour legend is a caption under the graph. Only produced when comparing at
least two traces exposing the NUMA topology.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
The lack of column and alignment in the header made it harder to read.

Before:
System: server_x_y_z Bios 3.2 Kernel 8.10.1

After:
System:server_x_y_z Bios:3.2 Kernel:8.10.1

Signed-off-by: Erwan Velu <e.velu@criteo.com>
With large data like cpu frequencies, the legend box tends to collide
with the Y-AXIS making it unreadable.

Let's move this box a bit away to the left to avid this.

Signed-off-by: Erwan Velu <e.velu@criteo.com>
The "No samples found in <metric>, ignoring metric." messages
polluted the output when a metric is missing.

A typical output was showing a missing PDU metric which can be a normal
thing.

So let's only print them if "--verbose" is requested.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
The stats output printed one max line per power metric, including metrics
that had no data at all, e.g. "PDU max : 0.00 Watts in " with an empty bench
name. That 0.00 line looked like a buggy processing from hwgraph rather than
what it actually was: a simply missing metric.

Metrics are now only displayed if valid data is found.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
Reading a value off the smp_scaling line graphs was imprecise with only the
major Y ticks.

Add a single intermediate horizontal gridline between two Y
ticks (AutoMinorLocator(2)), a bit bolder than the default minor grid so it is
actually readable, but still lighter than the solid major gridlines.

Applies to the scaling line graphs (perf, perf_watt, watts, cpu_clock, cpu_ipc); the
delta heatmaps are unaffected.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
The per-trace legend of the smp_scaling graphs shows, for each trace,
"<name> [min; mean; stddev; max]". Each label was padded to its own length,
so traces with different name lengths produced a misaligned table where the
[min; mean; stddev; max] columns did not line up, making the values harder to
read and compare between traces.

Pad the trace name and the values to a width common to all the traces of the
graph (like the environment graphs already do), so the columns stay aligned
whatever the trace names.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
For every per-core CPU metric (Core frequency, Core IPC and CPU Core
power consumption), render an additional graph showing how that metric
is distributed across the cores at steady state. Each core's mean value
over the run becomes one data point; a violin shows the density and an
overlaid box shows the median (red), mean (green dashed), quartiles and
outliers.

Where the line graphs answer "how did the metric evolve over time?",
these answer "how uniform were the cores?" -- immediately exposing
stragglers, bimodal behaviour (e.g. boosted vs throttled cores) and the
spread that a single averaged curve hides.

Like the other per-core graphs they are rendered twice, once over all
the cores (all_cores) and once restricted to the cores pinned during the
job (pinned_cores), and land in the same per-metric directories next to
their line graphs. The Y axis is autoscaled rather than zero-based: a
distribution is unreadable squished against a zero baseline.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
Mirror the per-core steady-state distribution graphs on the SMP scaling
views. For each trace and each per-core metric (Core frequency, Core IPC
and CPU Core power consumption), one graph plots a violin + box per
scaling step (X = worker count, evenly spaced), showing how the
core-to-core distribution of that metric evolves as the sweep grows.

The scaling line graphs plot a single averaged value per step; these
expose the spread that average hides -- cores that start to diverge or
throttle only past a given worker count, bimodal frequency behaviour, or
a widening per-core power spread. The box reports median (red), mean
(green dashed), quartiles and outliers.

Like the other per-core scaling graphs they are rendered for all_cores
and, when the sweep pins cores, pinned_cores, landing in the cpu_clock,
cpu_ipc and cpu_core_power directories next to the matching line graphs.
The Y axis is autoscaled rather than zero-based: a distribution is
unreadable squished against a zero baseline.

Why violin graphs at both levels: the per-benchmark (steady-state) and
the per-scaling-step distributions answer two complementary questions
that averages and time-series curves both flatten.

  - The individual, steady-state violin answers "for this one operating
    point, how uniform is the hardware?". A tight body means every core
    behaves alike; a long tail or a second lobe exposes a straggler, an
    asymmetric NUMA/boost domain or a mis-pinned core -- something a
    single mean value, and even the time-series line graph, hides because
    they collapse every core into one number per instant.

  - The scaling violins put those same distributions side by side along
    the load axis, so the shape itself becomes the signal: you watch the
    spread widen or split exactly at the worker count where cores begin
    to contend for power/thermal budget, and you can tell a uniform
    slow-down (body drops as a block) from a divergence (body stretches
    or splits) -- a distinction a scaling line of per-step averages
    cannot make.

Together they let a reader first confirm a single point is healthy, then
follow how that health degrades under increasing load, using the same
visual vocabulary at both zoom levels.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
Add a distribution view for per-core CPU metrics (frequency, IPC, core
power) aggregated by NUMA domain, complementing the existing per-domain
line graphs and heatmaps:

- numa_distribution_graph: one violin + box per NUMA domain for a single
  benchmark job (steady-state), landing next to the matching line graph
  in the same all_numa/pinned_numa directories.

- render_numa_scaling_ridgelines: the SMP-scaling counterpart. A first
  grouped-violin attempt (one violin per domain per step, all in one
  axes) became unreadable past a handful of domains and steps. After
  reviewing sample renderings of several alternatives (median+IQR line,
  small multiples, spread heatmap, ridgeline), the ridgeline design was
  picked: one panel per scaling step -- every step, laid out as a grid
  so the figure grows in rows rather than becoming unreadable -- each
  panel a stacked density (ridgeline) per NUMA domain, preserving the
  full distribution shape (skew, bimodality) that a single averaged
  value would flatten. The grid's header/footer are reserved as a
  constant number of inches (not a fraction of the figure) so they keep
  the same size regardless of how many rows a long sweep needs.

Both are rendered for all_numa (every core of each domain) and, when
the sweep pins cores, pinned_numa.

Also mark which CPU package each NUMA domain belongs to: hwbench does
not record this directly, but same-package domains are always much
closer to each other in the NUMA distance matrix than cross-package
ones, so _numa_domains_by_package groups domains whose mutual distance
is below a fixed threshold (union-find), with no need to know the
socket count up front. Each scaling ridgeline panel gets a pale
background wash per package behind its ridges, with a
"CPU package <n>" legend, colour-matched to the washes, centered below
the whole grid.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
Add two views to the SMP scaling output that make a sweep's scaling
quality legible against a perfect-linear reference. Both anchor the ideal
on the first scaling step (slope = perf0 / workers0), so a run that keeps
its first step's per-worker throughput would sit exactly on the ideal.

1. On the raw perf line graph, overlay per trace a dotted "ideal linear
   performance projection" in the trace's own colour. The projection
   lines are appended to the end of the per-trace statistics legend under
   a centered "linear performance projection" sub-header (a separate box
   never aligned cleanly with the stats table), separated by a blank row.
   Only drawn when the sweep has more than 3 steps.

2. A new per-trace graph, "Performance scaling linearity deviation",
   rendered in the perf directory: per scaling step it plots
   (measured / ideal - 1) * 100, so 0% is perfectly linear and a step at
   80% of the ideal reads -20%. The signed area to the 0% baseline is
   filled red below (scaling loss) / green above (superlinear) and the
   worst step is annotated. It reuses the perf graph's X axis (start at
   0, major tick every 8 workers, minor every 4) and label, adds the host
   info header via get_system_title() like the other per-trace graphs,
   and normalises the Y range across all rendered traces so their
   deviations can be compared at a glance.

Where the scaling line graphs answer "how fast did it go?", these answer
"how much scaling did we leave on the table, and where?".

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
When environmental graphs are rendered, it's hard to navigate into it
as it's difficult to remember what benchmark is doing what.

This commit is about adding a table that describes the benchmarks
that has been run and their context.

Before rendering any graph, write environment/by_host/<trace>/
benchmarks_summary.txt: a plain-text table with one row per benchmark
(avx_0, avx_1, ...), ordered by job_number, describing what each one
exercises -- job, engine/module, variant (engine_module_parameter),
worker count, pinned cores and per-run duration. It is a quick key to
what the (many) per-benchmark environmental graphs actually tested.

The header carries the same host description as the graphs, with
System / Bios / Kernel split onto their own aligned lines and the
Processor line, plus a summary of every NUMA node and its cores.

Formatting for readability at a glance:
  - Workers, NUMA nodes and Duration columns are right-aligned.
  - Pinned cores use the condensed, individually-bracketed range blocks
    with digits padded to 3 columns, so ranges line up even with
    3-digit core numbers.
  - The NUMA-nodes-spanned column is bracketed with node numbers padded
    to 2 digits (e.g. "[ 0- 3]").
  - Duration carries its unit in the header ("Duration (s)") and an
    integer value.

Generated up front in render_traces() (not as a pool task) so the file
is available before the graphs, and only when environmental graphs are
enabled.

A sample output looks like :

	Benchmarks summary - CZ2D2W09F0
	Study performance scalability

	System    : CZ2D2W09F0 ProLiant DL365 Gen11
	Bios      : v3.12
	Kernel    : 6.18.38-crto1264.el10.x86_64
	Processor : 2x AMD EPYC 9845 160-Core Processor - 320 physical cores and 20 NUMA domains

	NUMA nodes (20):
	  NUMA  0 : [  0- 15], [320-335]
	  NUMA  1 : [ 16- 31], [336-351]
	  NUMA  2 : [ 32- 47], [352-367]
	  NUMA  3 : [ 48- 63], [368-383]
	  NUMA  4 : [ 64- 79], [384-399]
	  NUMA  5 : [ 80- 95], [400-415]
	  NUMA  6 : [ 96-111], [416-431]
	  NUMA  7 : [112-127], [432-447]
	  NUMA  8 : [128-143], [448-463]
	  NUMA  9 : [144-159], [464-479]
	  NUMA 10 : [160-175], [480-495]
	  NUMA 11 : [176-191], [496-511]
	  NUMA 12 : [192-207], [512-527]
	  NUMA 13 : [208-223], [528-543]
	  NUMA 14 : [224-239], [544-559]
	  NUMA 15 : [240-255], [560-575]
	  NUMA 16 : [256-271], [576-591]
	  NUMA 17 : [272-287], [592-607]
	  NUMA 18 : [288-303], [608-623]
	  NUMA 19 : [304-319], [624-639]

	Benchmark  Job  Engine / module  Variant          Workers  Pinned cores          NUMA nodes  Duration (s)
	---------  ---  ---------------  ---------------  -------  --------------------  ----------  ------------
	avx_0      avx  stressng / vnni  avx_vpaddb512         32  [  0- 15], [320-335]        [ 0]            60
	avx_1      avx  stressng / vnni  avx_vpdpbusd512       32  [  0- 15], [320-335]        [ 0]            60
	avx_2      avx  stressng / vnni  avx_vpdpwssd512       32  [  0- 15], [320-335]        [ 0]            60
	avx_3      avx  stressng / vnni  avx_vpaddb512         64  [  0- 31], [320-351]     [ 0- 1]            60
	avx_4      avx  stressng / vnni  avx_vpdpbusd512       64  [  0- 31], [320-351]     [ 0- 1]            60
	avx_5      avx  stressng / vnni  avx_vpdpwssd512       64  [  0- 31], [320-351]     [ 0- 1]            60
	avx_6      avx  stressng / vnni  avx_vpaddb512         96  [  0- 47], [320-367]     [ 0- 2]            60
	avx_7      avx  stressng / vnni  avx_vpdpbusd512       96  [  0- 47], [320-367]     [ 0- 2]            60
	avx_8      avx  stressng / vnni  avx_vpdpwssd512       96  [  0- 47], [320-367]     [ 0- 2]            60
	avx_9      avx  stressng / vnni  avx_vpaddb512        128  [  0- 63], [320-383]     [ 0- 3]            60

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
Label the peak (max) point of every SMP scaling line with its value,
drawn in that line's own colour and using the same human-readable
formatting as the Y axis, so the top value can be read straight off the
graph. Covers all scaling line graph types (perf, perf_watt, watts,
cpu_clock, cpu_ipc).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
"SMP" isn't really meaningful here — these graphs are not about symmetric
multiprocessing as such, they show how a benchmark's performance (and the
related power/clock/IPC metrics) scales as workers are added. "Performance
scaling" describes what we are actually looking for.

Rename everywhere it surfaces: the graph titles (perf, perf_watt, watts,
cpu_clock, cpu_ipc and the NUMA delta heatmaps), the console progress
messages, the --no-scaling help text, and the output directory
(smp_scaling/ -> scaling/).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
Add a "Performance per worker distribution" violin+box graph to the scaling
output. Where the scaling line graph plots a single aggregated value per step,
this new graph opens that number up and shows how the individual stress-ng
worker threads performed: one violin+box per scaling step, built from the
stressor's own per-instance results.

This makes the spread across working threads visible at a glance -- whether
every worker runs at the same rate or some lag behind (imbalance, throttling,
a slow NUMA domain, noisy neighbours). It mirrors the per-core distribution
graphs rendered under all_cores/pinned_cores, but is built from the stressor's
individual results rather than the monitoring, and lands in the perf/ directory
next to the performance line graph.

To feed it, stress-ng's per-instance metrics are extracted from its YAML output
(--yaml, emitted next to stdout) into the bench "detail" key as a per-worker
"bogo op/s" list (using bogo-ops-per-second-real-time). Parsing uses PyYAML,
added as a dependency. The graph only renders when a run exposes such detail, so
older results are unaffected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
Give the violin+box distribution graphs the same two-tier Y grid as the
scaling linearity-deviation graph (solid major lines plus a fainter dashed
midline between the major Y ticks). The grid stays on the Y axis only:
vertical lines would cut through the violins and hide their shape, which is
the whole point of these graphs.

The two per-worker/per-core scaling distribution graphs also get the same
legend placement as the linearity graph (lower left) and surface the trace
name alongside the median/mean key, so a per-trace graph identifies its host
at a glance.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
The perf/ directory mixed the performance line graph with the linearity
deviation and per-worker distribution graphs, which cluttered the listing.
Give each derived analysis its own subdirectory (perf/linearity_deviation
and perf/per_worker_distribution) so the line graph stands alone and the
listing reads clearly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
The source-file box is drawn below the axes, but savefig(bbox_inches="tight")
was only told about the legends via bbox_extra_artists. Under square=True
(set_box_aspect(1)) the box could fall outside the tight crop and disappear,
as seen on the per-worker distribution graph. Track the box artist and pass it
to bbox_extra_artists so it is always part of the saved bounding box.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
When several traces are compared (>=2), summarize the whole scaling study
relative to the first trace (the reference), in max_versus/. Two outputs are
produced once from the versus task:
  - benchmarks_summary.txt: a text table, one section per benchmark type plus
    a closing aggregate section.
  - performance_summary_scorecard.png: an executive-summary heatmap of the
    aggregate (traces x metrics), each cell colored vs the reference.

Logic behind the design:
  - Reference-relative: the first trace is always the reference (1.00x) and
    every metric is a ratio to it -- comparing heterogeneous hosts only makes
    sense against a common baseline.
  - Full load: values are taken at each trace's max-workers step, the one point
    every sweep reaches; keeps it to one row per trace.
  - Per physical core (Delta perf/core) separates "bigger machine" from "faster
    core": raw throughput mostly tracks core count.
  - Aggregate is the geometric mean of the ratios -- the correct average for
    ratios spanning very different magnitudes across stressors.
  - Linearity is the deviation from perfect linear scaling (0% = linear), same
    definition as the linearity-deviation graph.
  - CoV (per-worker homogeneity) comes from the stress-ng per-instance detail;
    lower = workers ran more uniformly.
  - Scorecard color is oriented so green always means "better than reference"
    (power and CoV inverted); absolute values are printed in each cell, with the
    system description and a per-metric legend below the graph.
  - A benchmark the reference did not run (e.g. AVX-512 on a Zen2 reference) is
    reported as N/A and excluded from the aggregate.

Lives in graph/versus.py (a cross-trace comparison belongs with the versus
code) and is dispatched from the versus branch of _collect_plot_tasks.

benchmarks_summary.txt (partial sample):

  Reference : Zen2
    (ref) Zen2               1x AMD EPYC 7502P 32-Core Processor -  32 cores /  8 NUMA   scaling 8->64 workers
          Zen4c              1x AMD EPYC 8534P 64-Core Processor -  64 cores /  8 NUMA   scaling 16->128 workers
          Zen5c              2x AMD EPYC 9845 160-Core Processor - 320 cores / 20 NUMA   scaling 32->640 workers
          ...

  ### cpu/float128  [Bogo op/s]
  Trace            Wrk     Perf   Δperf  Perf/core  Δperf/core  Power  Δpower  Perf/W  Clock  Δclock   IPC   ΔIPC    CoV  Linearity
  ---------------  ---  -------  ------  ---------  ----------  -----  ------  ------  -----  ------  ----  -----  -----  ---------
  Zen2*          64  123.69K   1.00x      3.87K       1.00x   180W   1.00x   1.00x  3058M   1.00x  1.34  1.00x   0.6%        -9%
  Zen4c         128  302.38K   2.44x      4.72K       1.22x   194W   1.08x   2.26x  2921M   0.96x  1.70  1.27x   1.7%        -6%
    ...

  ### Aggregate (geometric mean of full-load ratios across all benchmarks)
  Trace             Δperf  Δperf/core  Δpower  Perf/W  Δclock   ΔIPC   CoV  Linearity
  ---------------  ------  ----------  ------  ------  ------  -----  ----  ---------
  Zen2*          1.00x       1.00x   1.00x   1.00x   1.00x  1.00x  0.4%        -9%
  Zen4c          2.17x       1.09x   1.07x   2.03x   0.94x  1.08x  0.9%        -8%
    ...

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…s not dropped

The raw performance scaling graph zipped the per-trace results against a
fixed 5-element `colors` list:

    for trace_name, color_name, e_color in zip(aggregated_perfs[perf], colors, cycle(e_colors)):

`zip()` stops at the shortest iterable, so with more than 5 traces the 6th
(and beyond) was silently truncated: it was never plotted, and — since the
legend/stats table is built from the plotted artists — never appeared in
the legend either. The error colours were already `cycle()`d, so only the
line colours capped the trace count.

Wrapping `colors` in `cycle(...)` like `e_colors` lets the sweep render any
number of traces (colours repeat past 5). The other scaling graphs render
one file per trace and were unaffected, which is why the drop looked
isolated to this single graph.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The traces comparison (max_versus/benchmarks_summary.txt) reported every
trace against the reference only at full load. Add the same report per
scaling step: one scaling/summary/benchmarks_summary_<N>_workers.txt file per
worker count, comparing every trace to the first at that step, so the delta
can be followed as the sweep grows rather than only at the top.

To share the format, the table/aggregate builder is factored out of
write_scaling_comparison into _render_traces_comparison(), driven by a few
scope strings (heading suffix, column note, aggregate header, and the reason
a benchmark the reference lacks is excluded). The full-load report now calls
it too; its data output is byte-identical (only the generalised Linearity /
Wrk column legend wording changed, now accurate for both reports).

Metrics gain a per-step source: _trace_job_step_metrics() returns
{workers: metrics} for every step with linearity anchored on the first step
(so the first step reads 100%); _trace_job_metrics() is now just its
max-workers entry, keeping the single report and the scorecard unchanged. A
trace with no run at a given worker count is reported as N/A in that step's
file. Wired into the existing scaling-comparison task (>=2 traces).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Erwan Velu <e.velu@criteo.com>
@ErwanAliasr1
ErwanAliasr1 force-pushed the hwgraph-comparison-reports branch from bdccfbd to 4d639cc Compare July 16, 2026 10:32
Comment thread graph/versus.py
benches = sorted((b for b in info["bench"] if not b.skipped()), key=lambda b: b.workers())
if not benches:
continue
perf_key = info["metrics"][0][0]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same comment as PR #125 : we only take the first metric. Again, just saying which it is in the title could be enough.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants