Lab / / benchmark

A 16×16 window at 9.92 ns against 4.85 ns by hand

The shape motion search actually computes: 16×16 at 4.85 ns by hand against 9.92 ns from the provider, with the 8×8 comparison carrying a kernel-shape difference.

SIMD and intrinsics, Benchmarking, Measurement methodologyWeft
Machine
Apple M5 Max, arm64
Commit
UNKNOWN
16×16, hand strided idiom
4.85 ns
16×16, weft provider
9.92 ns
16×16, weft generic scalar nest
94.6 ns
8×8, hand strided idiom
1.59 ns
8×8, weft provider
6.43 ns
8×8, weft generic scalar nest
24.2 ns

Two misaligned windows inside a 512×512 frame, with strides not equal to width — the shape motion search actually computes, rather than the contiguous case a benchmark finds convenient.

Nhand strided idiomweft providerweft generic (scalar nest)delta over hand
8×81.59 ns6.43 ns24.2 ns4.8 ns
16×164.85 ns9.92 ns94.6 ns5.1 ns
Nanoseconds per call, lower is better. The source gives absolute times and a delta over hand; it states no ratio, so none is given here.

**Environment.** Apple M5 Max, arm64, AdvSimd 128-bit fixed; BenchmarkDotNet 0.15.8, in-process; 8×8 and 16×16 blocks inside a 512×512 frame with strides ≠ width. Date and commit UNKNOWN.

**Methodology.** MOSTLY_MATCHED. Same machine and harness, one call per block on both sides — but the hand 8×8 is fully unrolled with two accumulators while the provider's width-8 path loops with drain bookkeeping. Lower is better for Weft, so both provider columns are losses on this shape.

The deltas over hand are 4.8 and 5.1 ns, which looks like one constant until you decompose it. Raw-delegate-minus-hand is about 1.4 ns at 16×16, which is true dispatch overhead, but about 3.3 ns at 8×8, where roughly 2 ns of it is kernel shape. So "kernel work at parity" holds cleanly at 16×16 and approximately at 8×8. That is why this is MOSTLY_MATCHED rather than MATCHED. The remainder is dispatch — delegate invocation and an argument array — that the hand baseline never pays because RyuJIT inlines it into the caller outright.

A shape lesson learned the measured way

An early version called the flat idiom once per row, so 8-wide rows never reached `UABD` at all — it needs 16 bytes — and ran pure scalar at 26 ns, slower than the 16×16 case. Shape matters more than instruction choice, and this is where I learned it rather than where I argued it.