Lab / failure

An exact tie went to Apple's sgemm

A native twin priced exactly the same as the managed kernel, the tie was broken the wrong way, and 144 kernels on the M4 quietly moved to Apple's sgemm. The decode got slower.

Compilers & SIMD, Negative resultsWeft
Qwen decode, M4, int8 NEON kernel
8.77 ms/token
After the tie went to sgemm
9.04 ms/token

Weft prices each candidate kernel and picks the cheapest. When native twins arrived, one priced exactly the same as the managed kernel it shadowed, and the tie was settled after selection rather than before.

The result on the M4 was that 144 Q8_0 matrix-vector kernels in Qwen's decode moved from the int8 NEON kernel to Apple's sgemm. Decode went from 8.77 to 9.04 ms per token, and the logits digest changed.

An exact tie now keeps the managed kernel. That rule came from this measurement, not from the design.