Skip to content

Audit the instructions in each build tarball against its CPU target - #318

Draft
HaoZeke wants to merge 18 commits into
EESSI:mainfrom
HaoZeke:tarball-isa-audit
Draft

HaoZeke wants to merge 18 commits into
EESSI:mainfrom
HaoZeke:tarball-isa-audit

Conversation

@HaoZeke

@HaoZeke HaoZeke commented Oct 11, 2026

Copy link
Copy Markdown

Depends on #311, which adds the per-target references this audit reads.
Until #311 is merged, the diff here also shows its commits; the last five commits are the new part.

#311 checks the compiler's native flags before a build.
This adds the matching check after the build: every instruction in the tarball has to be one the target CPU can run, against the same references.
A build for target T made on a node of a newer type passes every test on that node, and the result crashes with SIGILL on a real T machine.
The audit catches that before the tarball is deployed, however the build was made.

How it works.

  • scripts/isa_audit/check_tarball.sh unpacks the tarball and picks the newest GCC reference for the target.
  • x86_64: isa_audit.py decodes every executable section of every ELF file with iced-x86.
    An instruction fails when it needs a feature the reference marks -mno-X.
  • aarch64: isa_audit_aarch64.py disassembles each file with llvm-objdump twice, once with every feature and once with the reference's -mcpu.
    An instruction that decodes only in the first run fails.
  • Generic targets use scripts/isa_audit/references/{x86_64,aarch64}/generic/baseline.txt: what GCC enables for -march=x86-64 and -mcpu=generic, the flags EasyBuild's --optarch=GENERIC passes.
    generate_generic_references.sh writes them from a given GCC.
  • dispatch-allow.txt lists code that picks its instruction set at run time.
    Each line has a package and its features, optionally with function globs, and a comment naming the mechanism.
    Those sites print as DISPATCH and do not fail the build.
  • A violation prints an ERROR: line, which bot/check-build.sh already reports as a failed build.
  • Targets without a reference and hosts without iced-x86 or llvm-objdump are skipped with a message.
    EESSI_SKIP_ISA_AUDIT skips the audit.

Measured on 2025.06 with the gcc-14.3.0 references:

files against result
sapphirerapids hwloc 2.12.1 icelake fail: AVX512-FP16 vmovw in libhwloc.so
sapphirerapids GROMACS 2025.4 icelake fail: vmovw in all four libgromacs variants
icelake hwloc, icelake GROMACS, zen4 GROMACS their own references pass
whole intel/icelake tree, 1,263 package versions icelake 42 flagged, all runtime dispatch
whole intel/sapphirerapids tree icelake 402 flagged (32%), mostly AVX512-FP16
libzstd from nvidia/grace and neoverse_v1 neoverse_n1 fail: SVE (whilelo, ld1w)
libzstd from each Arm tree its own reference pass

A whole x86_64 tree takes about 25 minutes on 12 cores; one tarball from a PR takes seconds.

Generic targets.
I audited 25 x86_64/generic and 19 aarch64/generic installs from 2025.06 (BLAS, MPI, FFTW, GROMACS, SciPy-bundle, FFmpeg, codecs, compression, GCCcore's runtime libraries).
Most of what they flag is run-time dispatch, now on the list with function globs where the symbols allow it.
Two x86_64 installs fail for real:

  • UCX 1.19.0 is built with -mavx throughout: AVX shows up in 15 of its 16 ELF files.
    --enable-optimizations makes UCX's configure compile and run an AVX test program, and add -mavx for the whole package when it passes (checking -mavx... yes in the build log), so the generic install needs AVX.
  • SciPy-bundle 2025.07 has SSE3 (movddup, fisttp) in NumPy's baseline code, since NumPy's cpu-baseline=min is SSE3 on x86-64 (Enabled: SSE SSE2 SSE3 in the build log).
    The same code reaches scipy.special through npymath.

Neither is on the allow list, since neither is chosen at run time.
Three instruction patterns pass by rule because they are safe on any CPU of the architecture:

  • tzcnt is rep bsf, which a CPU without BMI1 runs as bsf; GCC emits it for __builtin_ctz at -march=x86-64.
  • xgetbv faults without OSXSAVE, so detection code runs it only after that CPUID check.
  • The -moutline-atomics helpers libgcc puts in each library use LSE only after checking __aarch64_have_lse_atomics; stripped libraries lose the helper names, so the auditor recognises the check itself.

Dispatch list.
The seed list came from the icelake tree and has 21 packages; the generic installs added OpenBLAS, OpenMPI, Boost, zstd, libjpeg-turbo, libwebp, pixman and more lines for GCCcore, x264, FFmpeg, libfabric and SciPy-bundle.
Each new line names the selection mechanism, and most carry function globs, e.g. ompi_op_avx_* for OpenMPI or *_bmi2* for zstd, so the same feature anywhere else in the package still fails.
OpenBLAS, FFmpeg and NumPy stay package-wide: FFmpeg is stripped, and the other two dispatch whole objects whose helpers have no target in their names.
The two FMA4 sites in Rust's rust-analyzer are compiler_builtins' libm fma_with_fma4, picked after a cpuid check, so any package with Rust code that calls f64::mul_add carries them.
The shadow-stack instructions from libgcc's unwinder (rdssp, then incssp only after a non-zero rdssp) pass by rule rather than per package.

Tests.
tests/isa_audit/test_check_tarball.sh compiles a small library, packs it as a tarball and checks it against native references, the generic baselines, allow lines with and without function globs and the outline-atomics helpers, stripped or not.
The rest cover a missing reference, EESSI_SKIP_ISA_AUDIT and a missing tarball.
All 21 pass locally with LLVM 19; the new workflow runs them on ubuntu-24.04 with LLVM 20.

Limitations.

  • Static archives (.a) are not read.
    Their code is audited in the package that links them.
  • Function globs need symbols; a stripped package can only have a package-wide line.
  • The bot hosts need iced-x86 and pyelftools for x86_64 and llvm-objdump for aarch64.
    Without them the audit skips with a message instead of failing.

Questions:

  1. Should a violation fail the build from the start, or print a warning until the allow list settles?
  2. iced-x86 comes from PyPI.
    Is a pip install on the bot hosts fine, or should it come from an EESSI module?
  3. UCX and NumPy above would fail a generic rebuild today.
    Should the fix go in their easyconfigs (no --enable-optimizations for UCX, cpu-baseline=none for NumPy), or should generic accept SSE3?

casparvl and others added 18 commits October 1, 2026 14:30
…building

To allow multiple (redundant) build hosts per CPU target, verify early in
EESSI-install-software.sh that the compilers of all toolchains supported in
the EESSI version being built translate the native architecture flag
(-march=native on x86_64, -mcpu=native on aarch64, as used by EasyBuild)
into exactly the same target flags as recorded in a reference for the CPU
target. This asks the GCC/Clang driver directly (via -###) rather than
relying on a surrogate such as lscpu.

- scripts/native_flags/get_supported_toolchains.py: extract the supported
  toplevel toolchains for an EESSI version from eb_hooks.py
- scripts/native_flags/check_native_flags.sh: check (or --generate)
  references in scripts/native_flags/references/<subdir>/<compiler>-<version>.txt
- references for all non-generic CPU targets built on the AWS build cluster

A missing reference is an error, unless
$EESSI_NATIVE_FLAGS_ALLOW_MISSING_REFERENCE is set. The check can be
skipped entirely with $EESSI_SKIP_NATIVE_FLAGS_CHECK, and is skipped for
generic builds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Caspar van Leeuwen <33718780+casparvl@users.noreply.github.com>
…SON file

The supported toplevel toolchains per EESSI version are now defined in
eessi_supported_toolchains.json rather than in eb_hooks.py itself, so that
they can also be used by other scripts without having to parse (or import)
the hooks file.

- eessi_supported_toolchains.json is located next to eb_hooks.py, both in
  the repository and when installed in <prefix>/init/easybuild/ by
  install_scripts.sh. eb_hooks.py locates it relative to its own location.
- Toolchains that can only be installed with a recent enough EasyBuild
  version (lfoss/2025b, rompi/2025a) now specify 'min_easybuild_version'
  instead of being appended conditionally in eb_hooks.py.
- CI also checks that the deployed eessi_supported_toolchains.json is
  up-to-date, like is done for eb_hooks.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Now that the supported toplevel toolchains are defined in
eessi_supported_toolchains.json, read them from there instead of extracting
them from eb_hooks.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Allow setting the location of the supported toolchains file through
  EESSI_SUPPORTED_TOOLCHAINS_FILE (default: next to eb_hooks.py), with
  specific errors for a missing file and for invalid JSON
- Document return value of load_supported_top_level_toolchains()
- Rename hook_files to easybuild_init_files in install_scripts.sh
- Rename test-eb-hooks.yml to test-eb-init-files.yml and check all init
  files in a single loop instead of duplicating the step
- Add unit tests for the JSON format and loader, run in CI only when
  relevant files change

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Only the structure of the shipped eessi_supported_toolchains.json is
tested now; the behaviour of load_supported_top_level_toolchains() is
tested using JSON files created on the fly, so the toolchain content is
not duplicated in the tests.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
eessi_supported_toolchains.json was converted to TOML in
separate-supported-toolchains, so parse that instead. Like eb_hooks.py,
respect $EESSI_SUPPORTED_TOOLCHAINS_FILE if set.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The native flags check runs before a build. Add the matching check after
the build: every instruction in the tarball has to be one the target can
run, against the same per-target references.

scripts/isa_audit/check_tarball.sh extracts the tarball and picks the
newest GCC reference for the target. x86_64 files are decoded with
iced-x86, and an instruction fails when it needs a feature the reference
marks -mno-X. aarch64 files are disassembled with llvm-objdump twice,
with all features and with the reference's features, and an instruction
that decodes only in the first fails. dispatch-allow.txt lists packages
that pick their instruction set at run time, and those are reported as
DISPATCH. A violation prints an ERROR line, which check-build.sh reports
as a failed build. Generic targets, targets without a reference and hosts
without iced-x86 or llvm-objdump are skipped with a message, and
EESSI_SKIP_ISA_AUDIT skips the audit.

bot/build.sh runs it after create_tarball.sh.
tests/isa_audit/test_check_tarball.sh compiles one loop into a shared
library, packs it the way create_tarball.sh lays out a tarball, and runs
check_tarball.sh on it. Built with -march=skylake-avx512 and checked
against haswell, it has to fail with an avx512f VIOLATION, and pass as
DISPATCH once an allow file lists the package. Built for haswell, it has
to pass against haswell. The aarch64 cases build with SVE and check it
against neoverse_n1 (fail) and neoverse_v1 (pass). The rest cover the
generic skip, a target without a reference, EESSI_SKIP_ISA_AUDIT and a
missing tarball.

The workflow installs gcc, the aarch64 cross compiler, LLVM 20, iced-x86
and pyelftools, and runs the script when the audit, the references or
the tests change.
The two vfmaddsd sites in bin/rust-analyzer sit in
compiler_builtins::math::libm_math::arch::x86::fma::fma_with_fma4, the
libm fma that compiler_builtins picks after a cpuid check. nm on the
2025.06 intel/icelake Rust 1.91.1 shows the symbol. Any Rust program that
calls f64::mul_add without FMA in its target features links the same
function, so packages with Rust code need fma4 on their line too.
Generic targets were skipped because EESSI#311 has no reference for them.
generate_generic_references.sh records what GCC enables for
-march=x86-64 and -mcpu=generic, the flags EasyBuild passes for optarch
GENERIC, in references/ next to the auditor, and
check_tarball.sh now uses them for x86_64/generic and aarch64/generic.

Against a baseline, a few instructions show up in code that is safe on
every CPU of the architecture, so they pass by rule:

- tzcnt is rep bsf, which runs as bsf without BMI1; GCC emits it for
  __builtin_ctz at -march=x86-64.
- xgetbv faults without OSXSAVE, so detection code runs it only after
  that CPUID check.
- libgcc's -moutline-atomics helpers use LSE only when
  __aarch64_have_lse_atomics is set; in stripped libraries they are
  found by that check instead of by name.

Allow lines can now end in function globs, so a line covers only the
dispatched functions and the same feature elsewhere in the package still
fails.
The aarch64 auditor takes the allow file as well, names the function
around each violation, and streams llvm-objdump output: holding the
disassembly of the aarch64 OpenBLAS in memory got the process killed.
Auditing 25 x86_64/generic and 19 aarch64/generic installs from 2025.06
against the new baselines flagged dispatch code in OpenBLAS, OpenMPI,
Boost, zstd, libjpeg-turbo, libwebp, pixman, x264, FFmpeg, libfabric,
NumPy and GCCcore's runtime libraries.
Each line names how the code is chosen and, where the symbols allow it,
lists the functions too, e.g. ompi_op_avx_* for OpenMPI or *_bmi2* for
zstd.
The old package-wide GCCcore line is split up the same way.
OpenBLAS, FFmpeg and NumPy stay package-wide: FFmpeg is stripped, and
the other two compile whole objects per target, with helpers that carry
no target in their names.

Two installs still fail and stay off the list, since nothing picks their
code at run time: UCX 1.19.0 is built with -mavx because its easyconfig
enables UCX's optimizations, and NumPy's baseline in SciPy-bundle
2025.07 includes SSE3.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants