Skip to content

AArch64 NRM2: recover finite results near overflow - #6095

Draft
BenKnill wants to merge 2 commits into
OpenMathLib:developfrom
BenKnill:fix/aarch64-nrm2-finite-overflow
Draft

BenKnill wants to merge 2 commits into
OpenMathLib:developfrom
BenKnill:fix/aarch64-nrm2-finite-overflow

Conversation

@BenKnill

@BenKnill BenKnill commented Oct 5, 2026

Copy link
Copy Markdown

Candidate repair for #6094. Adds an exact recovery path for large finite inputs to distinguish avoidable infinity from genuine overflow, retaining the existing ordinary-input arithmetic.

Tested in ARMV8 and NEOVERSEN1 configurations on native AArch64 Ubuntu using the 0.3.32 source tree. The same 12,148-case suite passed through CBLAS and Fortran in each configuration; existing unit tests also passed. Upstream regression-test integration, a full current-upstream build, and recovery-path timing remain outstanding.

The change covers the shared assembly kernels; zero-stride shortcuts, separate C kernels, and repeated-infinity behavior remain outside its scope.

Opened as a draft for discussion.

Opus 5.5 discovered this bug. GPT-6 Astra and GPT-6 Sol investigated and characterized the failure, then wrote and tested the candidate patch.

Guard the shared AArch64 norm kernels before their final multiplication and recover large finite inputs with exact integer sums of squares. Distinguish finite results from genuine overflow while retaining the ordinary arithmetic path.
@martin-frbg

Copy link
Copy Markdown
Collaborator

Thanks - the nrm2_range recovery should live with the other arm64-specific files in kernel/arm64 in my opinion (driver/others is for general setup code), and ultimately I think this would probably be solved by adopting the more thorough NRM2 scaling introduced by the recent Reference BLAS - which unfortunately requires rewriting kernels across all architectures. (But I haven't run your reproducer against said Reference BLAS yet)

Move nrm2_range.c from driver/others to kernel/arm64, next to the nrm2.S
and znrm2.S kernels that call it, and restore the driver/others build
files. The helper is built once per library as an arm64-only common
kernel object, in the same way as x86's kernel/x86/cpuid.S, so
DYNAMIC_ARCH builds still contain a single copy. No code changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@BenKnill

BenKnill commented Oct 5, 2026

Copy link
Copy Markdown
Author

Thanks. Moved: nrm2_range.c is now in kernel/arm64, built once per library as an arm64-only common kernel object, as x86 builds cpuid.S. Its arm64-guarded build lines are in kernel/Makefile and kernel/CMakeLists.txt (the KERNEL.* files only set variables), and the PR no longer touches driver/others. For ARMV8 and NEOVERSEN1 the 12,148-case suite output is unchanged, with no spurious infinities in the tested cases.

Against Reference BLAS at LAPACK master 6de9594, the reproducer returns finite results in round-to-nearest, but its double-precision case returns inf under FE_UPWARD. A targeted search also found longer vectors whose exact norm is below DBL_MAX but which return inf in round-to-nearest. For example, this eight-element vector, about 0.05 ulp below DBL_MAX, returns inf on both arm64 and x86-64:

0x1.bd2788e9e8b9ap+1023 0x1.cb33a7ccd84f4p+1021 0x1.52c1bd33cfe1fp+1021 0x1.fb2a0fcb66de1p+1021
0x1.6c278b240b52dp+1021 0x1.1f0c45e85aa4dp+1021 0x1.27cf1d741ec94p+1021 0x1.73f734386dff4p+1021

These cases suggest that adopting the current Reference BLAS scaling would still leave some avoidable-overflow cases. The candidate recovery handles them in the tested arm64 builds.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants