Lab / / experiment

A flat result, reported flat

One, two and four accumulator chains measured 356.3 / 356.2 / 356.6 ns. The dependency-chain argument does not apply to this kernel shape.

SIMD and intrinsics, Negative results, BenchmarkingWeft
Machine
Apple M5 Max, arm64
Commit
UNKNOWN
1 accumulator chain
356.3 ns
2 accumulator chains
356.2 ns
4 accumulator chains
356.6 ns

I implemented multi-accumulator chains for the SAD kernel because the dependency-chain argument said I should. Then I measured it. 356.3 ns at one chain, 356.2 ns at two, 356.6 ns at four.

That is flat to within noise. Lower is better here, and there is no lower to be had.

Why the argument did not apply

The dependency-chain argument does not describe this kernel shape. The kernel is bound by the twelve-op widening chain, not by the accumulator's latency, so splitting the accumulator gives the out-of-order engine nothing new to work with. The machinery is implemented and guarded correctly. On this workload it is simply not the bottleneck.

I kept it as a negative result rather than deleting the code path or quietly dropping the experiment from the report. A benchmark that is not being used as an advertisement is supposed to look like this.

Accumulator chainsns per call
1 (baseline)356.3
2356.2
4356.6
Direction: lower is better for all three arms. The spread is the noise floor, not an effect.