Lab / / benchmark
A 16×16 window at 9.92 ns against 4.85 ns by hand
The shape motion search actually computes: 16×16 at 4.85 ns by hand against 9.92 ns from the provider, with the 8×8 comparison carrying a kernel-shape difference.
SIMD and intrinsics, Benchmarking, Measurement methodologyWeft- Machine
- Apple M5 Max, arm64
- Commit
- UNKNOWN
- 16×16, hand strided idiom
- 4.85 ns
- 16×16, weft provider
- 9.92 ns
- 16×16, weft generic scalar nest
- 94.6 ns
- 8×8, hand strided idiom
- 1.59 ns
- 8×8, weft provider
- 6.43 ns
- 8×8, weft generic scalar nest
- 24.2 ns
Two misaligned windows inside a 512×512 frame, with strides not equal to width — the shape motion search actually computes, rather than the contiguous case a benchmark finds convenient.
| N | hand strided idiom | weft provider | weft generic (scalar nest) | delta over hand |
|---|---|---|---|---|
| 8×8 | 1.59 ns | 6.43 ns | 24.2 ns | 4.8 ns |
| 16×16 | 4.85 ns | 9.92 ns | 94.6 ns | 5.1 ns |
**Environment.** Apple M5 Max, arm64, AdvSimd 128-bit fixed; BenchmarkDotNet 0.15.8, in-process; 8×8 and 16×16 blocks inside a 512×512 frame with strides ≠ width. Date and commit UNKNOWN.
**Methodology.** MOSTLY_MATCHED. Same machine and harness, one call per block on both sides — but the hand 8×8 is fully unrolled with two accumulators while the provider's width-8 path loops with drain bookkeeping. Lower is better for Weft, so both provider columns are losses on this shape.
The deltas over hand are 4.8 and 5.1 ns, which looks like one constant until you decompose it. Raw-delegate-minus-hand is about 1.4 ns at 16×16, which is true dispatch overhead, but about 3.3 ns at 8×8, where roughly 2 ns of it is kernel shape. So "kernel work at parity" holds cleanly at 16×16 and approximately at 8×8. That is why this is MOSTLY_MATCHED rather than MATCHED. The remainder is dispatch — delegate invocation and an argument array — that the hand baseline never pays because RyuJIT inlines it into the caller outright.
A shape lesson learned the measured way
An early version called the flat idiom once per row, so 8-wide rows never reached `UABD` at all — it needs 16 bytes — and ran pure scalar at 26 ns, slower than the 16×16 case. Shape matters more than instruction choice, and this is where I learned it rather than where I argued it.