Conversation
Merging this PR will regress 1 benchmark
|
| Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|
| ❌ | xshape_initializer[std\:\:array<std\:\:size_t, 4>] |
151.7 ns | 180.8 ns | -16.13% |
| ⚡ | xshape_access[std\:\:array<std\:\:size_t, 4>] |
154.4 ns | 125.3 ns | +23.28% |
| 🆕 | broadcast_dynamic_raw |
N/A | 594.7 µs | N/A |
| 🆕 | broadcast_dynamic_xtensor |
N/A | 1.2 ms | N/A |
| 🆕 | linear_dynamic_raw |
N/A | 1.1 ms | N/A |
| 🆕 | linear_dynamic_xtensor |
N/A | 1.4 ms | N/A |
| 🆕 | linear_fixed_xtensor |
N/A | 4.8 µs | N/A |
| 🆕 | transpose_cast_raw |
N/A | 400.7 µs | N/A |
| 🆕 | transpose_cast_xtensor |
N/A | 456.8 µs | N/A |
create_xview |
< 1 ns | < 1 ns | N/A |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing wolfv:perf-upstream-master (68f5541) with master (d9a57b6)
|
CodSpeed’s remaining failure is unrelated measurement noise: this PR does not touch |
Summary
This PR improves assignment performance for dynamic-rank broadcast and permutation expressions, focusing on the stepper-heavy workloads reported in #2784, #2683, and #2859.
It adds two conservative execution paths:
Unsupported types, dimensions, layouts, views, and stride patterns retain the existing assignment path. The PR also adds comparison benchmarks for raw loops, Eigen, Armadillo (when available), dynamic shapes, fixed shapes, broadcasting, and transpose/cast expressions.
Runtime broadcast plan
For non-trivial three-dimensional row-major broadcasts with directly addressable arithmetic leaves:
This avoids rebuilding/reseeking recursive steppers for every short inner row.
Permutation kernels
For dense 3-D
cast(transpose(...)) / scalarassignments:Benchmarks
benchmark_compare.cppcovers dynamic contiguous fused arithmetic, dynamic N-D broadcasting, fixed-shape fused arithmetic, transpose/cast/scaling, raw loops, Eigen Array/Tensor, and optional Armadillo comparisons.Measured with
-O3 -march=native, xsimd 14, AVX-512, SIMD width 8:{64,64,16}{128,256,3} -> {3,128,256}Validation
xtest).