Lab / / experiment
A flat result, reported flat
One, two and four accumulator chains measured 356.3 / 356.2 / 356.6 ns. The dependency-chain argument does not apply to this kernel shape.
SIMD and intrinsics, Negative results, BenchmarkingWeft- Machine
- Apple M5 Max, arm64
- Commit
- UNKNOWN
- 1 accumulator chain
- 356.3 ns
- 2 accumulator chains
- 356.2 ns
- 4 accumulator chains
- 356.6 ns
I implemented multi-accumulator chains for the SAD kernel because the dependency-chain argument said I should. Then I measured it. 356.3 ns at one chain, 356.2 ns at two, 356.6 ns at four.
That is flat to within noise. Lower is better here, and there is no lower to be had.
Why the argument did not apply
The dependency-chain argument does not describe this kernel shape. The kernel is bound by the twelve-op widening chain, not by the accumulator's latency, so splitting the accumulator gives the out-of-order engine nothing new to work with. The machinery is implemented and guarded correctly. On this workload it is simply not the bottleneck.
I kept it as a negative result rather than deleting the code path or quietly dropping the experiment from the report. A benchmark that is not being used as an advertisement is supposed to look like this.
| Accumulator chains | ns per call |
|---|---|
| 1 (baseline) | 356.3 |
| 2 | 356.2 |
| 4 | 356.6 |