Lab / / failure

24 bytes inside the timed region

An earlier SAD run reported the idiom at 15.16 ns instead of 4.24 ns. The cause was 24 bytes of allocation per operation inside the measured region.

Measurement methodology, Benchmarking, Negative resultsWeft
Machine
Apple M5 Max, arm64
Commit
UNKNOWN
Corrected, N=256
4.24 ns
Earlier invalid run, N=256
15.16 ns
Distortion
3.6×
Allocated per operation, inside the timed region
24 bytes

An earlier run of the SAD benchmark reported the idiom at 15.16 ns for N=256. The corrected figure is 4.24 ns. That is a 3.6× distortion, and it came from 24 bytes allocated per operation inside the timed region.

The allocation was `ExecutionBinding.Output(name)`, which used a LINQ `FirstOrDefault` and allocated an enumerator on every call. Nothing was wrong with the kernel. The harness was charging its own bookkeeping to the code under test.

Why it hurt the smallest case most

Harness overhead is close to fixed, so it is proportionally worst where the kernel is shortest. N=256 is the smallest input in the set, and it is the case the artifact damaged most. The lookup is now allocation-free and the benchmark reports zero allocations on every row.

The report names this the fourth such artifact in M2, after `stackalloc` zeroing, a loop-versus-unroll confound, and overhead correlated with the independent variable. Four methodology bugs found and fixed says more about the rigor of the file than any single number in it does.

RunN=256 readingStatus
earlier15.16 nsinvalid — 24 B/op allocated in the timed region
corrected4.24 nsallocations reported as zero on every row
Direction: the lower figure is the correct one; the higher figure was produced by the harness.