Repository navigation
Conversation
…building To allow multiple (redundant) build hosts per CPU target, verify early in EESSI-install-software.sh that the compilers of all toolchains supported in the EESSI version being built translate the native architecture flag (-march=native on x86_64, -mcpu=native on aarch64, as used by EasyBuild) into exactly the same target flags as recorded in a reference for the CPU target. This asks the GCC/Clang driver directly (via -###) rather than relying on a surrogate such as lscpu. - scripts/native_flags/get_supported_toolchains.py: extract the supported toplevel toolchains for an EESSI version from eb_hooks.py - scripts/native_flags/check_native_flags.sh: check (or --generate) references in scripts/native_flags/references/<subdir>/<compiler>-<version>.txt - references for all non-generic CPU targets built on the AWS build cluster A missing reference is an error, unless $EESSI_NATIVE_FLAGS_ALLOW_MISSING_REFERENCE is set. The check can be skipped entirely with $EESSI_SKIP_NATIVE_FLAGS_CHECK, and is skipped for generic builds. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Caspar van Leeuwen <33718780+casparvl@users.noreply.github.com>
…SON file The supported toplevel toolchains per EESSI version are now defined in eessi_supported_toolchains.json rather than in eb_hooks.py itself, so that they can also be used by other scripts without having to parse (or import) the hooks file. - eessi_supported_toolchains.json is located next to eb_hooks.py, both in the repository and when installed in <prefix>/init/easybuild/ by install_scripts.sh. eb_hooks.py locates it relative to its own location. - Toolchains that can only be installed with a recent enough EasyBuild version (lfoss/2025b, rompi/2025a) now specify 'min_easybuild_version' instead of being appended conditionally in eb_hooks.py. - CI also checks that the deployed eessi_supported_toolchains.json is up-to-date, like is done for eb_hooks.py. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Now that the supported toplevel toolchains are defined in eessi_supported_toolchains.json, read them from there instead of extracting them from eb_hooks.py. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Allow setting the location of the supported toolchains file through EESSI_SUPPORTED_TOOLCHAINS_FILE (default: next to eb_hooks.py), with specific errors for a missing file and for invalid JSON - Document return value of load_supported_top_level_toolchains() - Rename hook_files to easybuild_init_files in install_scripts.sh - Rename test-eb-hooks.yml to test-eb-init-files.yml and check all init files in a single loop instead of duplicating the step - Add unit tests for the JSON format and loader, run in CI only when relevant files change Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Only the structure of the shipped eessi_supported_toolchains.json is tested now; the behaviour of load_supported_top_level_toolchains() is tested using JSON files created on the fly, so the toolchain content is not duplicated in the tests. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
eessi_supported_toolchains.json was converted to TOML in separate-supported-toolchains, so parse that instead. Like eb_hooks.py, respect $EESSI_SUPPORTED_TOOLCHAINS_FILE if set. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The native flags check runs before a build. Add the matching check after the build: every instruction in the tarball has to be one the target can run, against the same per-target references. scripts/isa_audit/check_tarball.sh extracts the tarball and picks the newest GCC reference for the target. x86_64 files are decoded with iced-x86, and an instruction fails when it needs a feature the reference marks -mno-X. aarch64 files are disassembled with llvm-objdump twice, with all features and with the reference's features, and an instruction that decodes only in the first fails. dispatch-allow.txt lists packages that pick their instruction set at run time, and those are reported as DISPATCH. A violation prints an ERROR line, which check-build.sh reports as a failed build. Generic targets, targets without a reference and hosts without iced-x86 or llvm-objdump are skipped with a message, and EESSI_SKIP_ISA_AUDIT skips the audit. bot/build.sh runs it after create_tarball.sh.
tests/isa_audit/test_check_tarball.sh compiles one loop into a shared library, packs it the way create_tarball.sh lays out a tarball, and runs check_tarball.sh on it. Built with -march=skylake-avx512 and checked against haswell, it has to fail with an avx512f VIOLATION, and pass as DISPATCH once an allow file lists the package. Built for haswell, it has to pass against haswell. The aarch64 cases build with SVE and check it against neoverse_n1 (fail) and neoverse_v1 (pass). The rest cover the generic skip, a target without a reference, EESSI_SKIP_ISA_AUDIT and a missing tarball. The workflow installs gcc, the aarch64 cross compiler, LLVM 20, iced-x86 and pyelftools, and runs the script when the audit, the references or the tests change.
The two vfmaddsd sites in bin/rust-analyzer sit in compiler_builtins::math::libm_math::arch::x86::fma::fma_with_fma4, the libm fma that compiler_builtins picks after a cpuid check. nm on the 2025.06 intel/icelake Rust 1.91.1 shows the symbol. Any Rust program that calls f64::mul_add without FMA in its target features links the same function, so packages with Rust code need fma4 on their line too.
Generic targets were skipped because EESSI#311 has no reference for them. generate_generic_references.sh records what GCC enables for -march=x86-64 and -mcpu=generic, the flags EasyBuild passes for optarch GENERIC, in references/ next to the auditor, and check_tarball.sh now uses them for x86_64/generic and aarch64/generic. Against a baseline, a few instructions show up in code that is safe on every CPU of the architecture, so they pass by rule: - tzcnt is rep bsf, which runs as bsf without BMI1; GCC emits it for __builtin_ctz at -march=x86-64. - xgetbv faults without OSXSAVE, so detection code runs it only after that CPUID check. - libgcc's -moutline-atomics helpers use LSE only when __aarch64_have_lse_atomics is set; in stripped libraries they are found by that check instead of by name. Allow lines can now end in function globs, so a line covers only the dispatched functions and the same feature elsewhere in the package still fails. The aarch64 auditor takes the allow file as well, names the function around each violation, and streams llvm-objdump output: holding the disassembly of the aarch64 OpenBLAS in memory got the process killed.
Auditing 25 x86_64/generic and 19 aarch64/generic installs from 2025.06 against the new baselines flagged dispatch code in OpenBLAS, OpenMPI, Boost, zstd, libjpeg-turbo, libwebp, pixman, x264, FFmpeg, libfabric, NumPy and GCCcore's runtime libraries. Each line names how the code is chosen and, where the symbols allow it, lists the functions too, e.g. ompi_op_avx_* for OpenMPI or *_bmi2* for zstd. The old package-wide GCCcore line is split up the same way. OpenBLAS, FFmpeg and NumPy stay package-wide: FFmpeg is stripped, and the other two compile whole objects per target, with helpers that carry no target in their names. Two installs still fail and stay off the list, since nothing picks their code at run time: UCX 1.19.0 is built with -mavx because its easyconfig enables UCX's optimizations, and NumPy's baseline in SciPy-bundle 2025.07 includes SSE3.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Depends on #311, which adds the per-target references this audit reads.
Until #311 is merged, the diff here also shows its commits; the last five commits are the new part.
#311 checks the compiler's native flags before a build.
This adds the matching check after the build: every instruction in the tarball has to be one the target CPU can run, against the same references.
A build for target T made on a node of a newer type passes every test on that node, and the result crashes with SIGILL on a real T machine.
The audit catches that before the tarball is deployed, however the build was made.
How it works.
scripts/isa_audit/check_tarball.shunpacks the tarball and picks the newest GCC reference for the target.isa_audit.pydecodes every executable section of every ELF file with iced-x86.An instruction fails when it needs a feature the reference marks
-mno-X.isa_audit_aarch64.pydisassembles each file withllvm-objdumptwice, once with every feature and once with the reference's-mcpu.An instruction that decodes only in the first run fails.
scripts/isa_audit/references/{x86_64,aarch64}/generic/baseline.txt: what GCC enables for-march=x86-64and-mcpu=generic, the flags EasyBuild's--optarch=GENERICpasses.generate_generic_references.shwrites them from a given GCC.dispatch-allow.txtlists code that picks its instruction set at run time.Each line has a package and its features, optionally with function globs, and a comment naming the mechanism.
Those sites print as DISPATCH and do not fail the build.
ERROR:line, whichbot/check-build.shalready reports as a failed build.llvm-objdumpare skipped with a message.EESSI_SKIP_ISA_AUDITskips the audit.Measured on 2025.06 with the gcc-14.3.0 references:
hwloc2.12.1vmovwinlibhwloc.sovmovwin all fourlibgromacsvariantshwloc, icelake GROMACS, zen4 GROMACSlibzstdfrom nvidia/grace and neoverse_v1whilelo,ld1w)libzstdfrom each Arm treeA whole x86_64 tree takes about 25 minutes on 12 cores; one tarball from a PR takes seconds.
Generic targets.
I audited 25 x86_64/generic and 19 aarch64/generic installs from 2025.06 (BLAS, MPI, FFTW, GROMACS, SciPy-bundle, FFmpeg, codecs, compression, GCCcore's runtime libraries).
Most of what they flag is run-time dispatch, now on the list with function globs where the symbols allow it.
Two x86_64 installs fail for real:
-mavxthroughout: AVX shows up in 15 of its 16 ELF files.--enable-optimizationsmakes UCX's configure compile and run an AVX test program, and add-mavxfor the whole package when it passes (checking -mavx... yesin the build log), so the generic install needs AVX.movddup,fisttp) in NumPy's baseline code, since NumPy'scpu-baseline=minis SSE3 on x86-64 (Enabled: SSE SSE2 SSE3in the build log).The same code reaches
scipy.specialthroughnpymath.Neither is on the allow list, since neither is chosen at run time.
Three instruction patterns pass by rule because they are safe on any CPU of the architecture:
tzcntisrep bsf, which a CPU without BMI1 runs asbsf; GCC emits it for__builtin_ctzat-march=x86-64.xgetbvfaults without OSXSAVE, so detection code runs it only after that CPUID check.-moutline-atomicshelpers libgcc puts in each library use LSE only after checking__aarch64_have_lse_atomics; stripped libraries lose the helper names, so the auditor recognises the check itself.Dispatch list.
The seed list came from the icelake tree and has 21 packages; the generic installs added OpenBLAS, OpenMPI, Boost, zstd, libjpeg-turbo, libwebp, pixman and more lines for GCCcore, x264, FFmpeg, libfabric and SciPy-bundle.
Each new line names the selection mechanism, and most carry function globs, e.g.
ompi_op_avx_*for OpenMPI or*_bmi2*for zstd, so the same feature anywhere else in the package still fails.OpenBLAS, FFmpeg and NumPy stay package-wide: FFmpeg is stripped, and the other two dispatch whole objects whose helpers have no target in their names.
The two FMA4 sites in Rust's
rust-analyzerarecompiler_builtins' libmfma_with_fma4, picked after a cpuid check, so any package with Rust code that callsf64::mul_addcarries them.The shadow-stack instructions from libgcc's unwinder (
rdssp, thenincssponly after a non-zerordssp) pass by rule rather than per package.Tests.
tests/isa_audit/test_check_tarball.shcompiles a small library, packs it as a tarball and checks it against native references, the generic baselines, allow lines with and without function globs and the outline-atomics helpers, stripped or not.The rest cover a missing reference,
EESSI_SKIP_ISA_AUDITand a missing tarball.All 21 pass locally with LLVM 19; the new workflow runs them on ubuntu-24.04 with LLVM 20.
Limitations.
.a) are not read.Their code is audited in the package that links them.
iced-x86andpyelftoolsfor x86_64 andllvm-objdumpfor aarch64.Without them the audit skips with a message instead of failing.
Questions:
iced-x86comes from PyPI.Is a pip install on the bot hosts fine, or should it come from an EESSI module?
Should the fix go in their easyconfigs (no
--enable-optimizationsfor UCX,cpu-baseline=nonefor NumPy), or should generic accept SSE3?