Lab / failure
An exact tie went to Apple's sgemm
A native twin priced exactly the same as the managed kernel, the tie was broken the wrong way, and 144 kernels on the M4 quietly moved to Apple's sgemm. The decode got slower.
Compilers & SIMD, Negative resultsWeft- Qwen decode, M4, int8 NEON kernel
- 8.77 ms/token
- After the tie went to sgemm
- 9.04 ms/token
Weft prices each candidate kernel and picks the cheapest. When native twins arrived, one priced exactly the same as the managed kernel it shadowed, and the tie was settled after selection rather than before.
The result on the M4 was that 144 Q8_0 matrix-vector kernels in Qwen's decode moved from the int8 NEON kernel to Apple's sgemm. Decode went from 8.77 to 9.04 ms per token, and the logits digest changed.
An exact tie now keeps the managed kernel. That rule came from this measurement, not from the design.