Fuse it only if it pays
Weft used to fuse attention wherever it recognised the shape. The new composition engine prices every region on the machine in front of it and races the close calls. The first thing it did was undo a fusion.
1 Oct 2026 · 7 min · Ernest
Compilers & SIMD, Negative resultsIn the last piece about Weft I said the win was in how kernels get put together, not in the kernels. What I didn't say was how the putting-together worked, because in hindsight it's a little embarrassing: pattern matchers. A pass walked the plan, spotted the shape of an attention chain (QKᵀ, a row max, exp, a row sum, a divide, PV), and swapped the six kernels for one fused body I'd written for exactly that shape. If the pattern matched, it fused. Whether fusing paid was not a question anyone asked.
It worked on the machine I wrote it on. The fused attention body cleared ggml's own attention by 1.26 to 1.32× per head on a Zen+ box, and when I switched the three x64 fusion gates on in early September, the whole SAM 3 image model (pixels in, masks out) lost 29% of its wall time on a Threadripper and came out 1.68× faster than sam3.cpp, the ggml port it's measured against, in a smoke run of four processes per arm. That's under Weft's Inference contract, the one that lets it reassociate float math. An online softmax reorders the sums, so under the Strict contract none of this fires and attention stays six kernels.
Same body, different machine
Then I measured the same body on a rented Broadwell Xeon, one core pinned, against ggml's flash attention at SAM 3's shape, 5,184 tokens by 32. ggml took 113 ms per head. My fused body took 175. The thing the matcher had been inserting on sight was 1.54× slower than the reference it was supposed to beat, and the matcher had no way to know, because it never priced anything.
The body itself was fixable. I rewrote it with the queries as the vector lanes instead of the keys: the softmax becomes lane-wise with no horizontal folds, K and V stay in their row layout instead of being transposed into 648 KiB of scratch, both contractions sit in twelve accumulators across the sixteen YMM registers with nothing spilled, and a query tile is sized from the measured L2 so K and V leave L3 eleven times per head instead of 1,296. Same operations in the same order, so the output is bit-identical to the old body. 64 ms per head: 1.76× faster than ggml instead of 1.54× slower. Whole model on the same box, 17.9 s down to 12.5 s, and against ggml's own run of the detector there, Weft went from 1.45× ahead to 2.07×.
A week later the three segmentation-head convolutions got the same treatment on a Zen 2 Ryzen: a blocked nest, no spilled accumulators, bit-exact under both contracts, 4 to 6× faster than ggml's im2col-plus-matmul per kernel. Which leaves the x64 number I was unhappy about last time looking like this.
| Machine | Contract | When | ggml ÷ Weft |
|---|---|---|---|
| x64, host not recorded (the previous article) | not recorded | undated | 0.27× (3.75× behind) |
| Xeon E5-2680 v4, Broadwell | Inference | 24 September | 2.07× |
| Ryzen 9 3900X, Zen 2 | Inference | 25 September | 2.29 to 2.31× |
| Ryzen 9 3900X, Zen 2 | Strict | 25 September | 1.04 to 1.05× |
Under Strict it's level, which is fair: Strict forbids the reassociation the online softmax depends on, so the biggest fusion in the model isn't available there. The language-model side is flatter. Weft's Qwen 2.5 decode is 1.2× ahead of llama.cpp at one thread on a Ryzen 5950X and about 10% behind it at four threads or more, where both sit at the memory roof.
A matcher can't price
Look at what those two sittings say. Per head the new attention body was 1.76× ggml, and in the whole model it bought 1.43×. The conv arm was 4 to 6× per kernel and about 1.25× on the wall. The number that matters is the one in context, and the context changes with the machine: a body admitted on a Zen+ win lost on Broadwell. Pattern matchers fuse the pattern. I wanted something that fuses the region only when the price, on this machine, in this plan, says so. That's the composition engine that merged on 1 October.
It has no patterns. A region starts from any kernel as its anchor, the one whose output the outside will see, and grows backwards through the producers of its inputs. A producer joins when everything that reads its output is already inside, so the intermediate becomes something nobody outside needs, and when the grown set still shares a common tile axis so it can be lowered as one nest. Every step of that growth is a candidate boundary. Attention falls out of it. So does a norm and its consumer, a contraction with its elementwise tail, and the two-segment decode step in Qwen, none of which anyone wrote a matcher for.
There is exactly one algebraic rewrite, and it's the only reason attention can be a single pass. Take a sum whose terms carry a factor that's constant along the fold axis but only known after a different fold over the same axis has finished: softmax's max, and then its total. Scaling distributes over addition, so you can fold against a running reference and multiply the correction in afterwards, one pass instead of three. The engine derives that from the member bodies rather than knowing what a softmax is, and it admits the rewrite only under Inference, because it reassociates. For Strict there's now a separate exact fused attention body, query-tiled, every fold in the oracle's order, bit-identical to the six kernels. It's offered to the same engine as a library body, and its walls aren't measured yet.
Deciding is three phases. First the cost model, in cycles: the members unfused, with their operands at the reuse distances the plan already gives them, against every tiling of the fused program the engine can lower, with branch-and-bound so a region that can't beat its unfused price stops being explored. Identical regions, every head of every block, are priced once. Second, where the model's margin is thin, a timed race inside the compile budget between the unfused members, the best fused program and any library body. Third, claim the winners and write the decision to a record, so a warm compile reuses it instead of racing again. On a plan the engine has decided the old matchers don't run at all. WEFT_GENERIC_FUSION=off puts them back, which is how the comparison below was done.
The first thing it did was un-fuse something
The one clean measurement so far is on an Apple M4: SAM 3 under Inference, four processes per arm, order reversed each round. The matchers' plan, with their fused attention: 2,501 to 2,522 ms. The engine's plan, which left that attention unfused: 1,715 to 1,722 ms. Every process of one arm is faster than every process of the other. On this machine the fusion the matchers had been applying on sight was a 1.46× loss, and the right call was to not do it. In a preliminary round where the engine also formed 432 regions of its own, the wall was 1,651 to 1,702 ms, a little better again. Those rounds ran without an A/A arm, so the numbers sit beside their within-arm spread (under 1%) rather than a measured floor. The driver carries the control now.
I'd like to tell you the search was clever. In the third sitting it searched for four seconds and decided nothing. One signature's precise price took 2,251 ms of a 2,000 ms half-budget (the budget was checked before starting a signature, and a started price couldn't be interrupted), everything behind it was skipped, 4 of 644 signatures got priced, and the 2,851 undecided anchors were handed back to the matchers. So the engine's plan came out byte-identical to the matchers' plan, the two arms measured the same thing, and the reducer printed ok on a ratio of 0.9999.
Two smaller things it taught me. A raced verdict is a timed comparison, so from the same plan on the same tree five of six processes chose 432 regions and one chose 384, which is why the decision record exists. And the plan hash is useless for comparing two searching compiles, because the search writes its own elapsed milliseconds into the plan's note. Compare the kernel set and the output digest instead.
What binds now is how fast the model can price: around 4 ms per kernel price, 644 signatures on SAM 3, two seconds to spend. The engine reaches a handful before the budget runs out and leaves the rest unfused, which happens to be the right answer on the M4 and is no answer at all elsewhere. Making the pricer cheap enough to reach the attention boundaries inside the budget is the current lane. The engine's own x64 and Qwen walls are still open, so the x64 rows in the table are the new bodies under the old gates. The matchers are still in the tree, demoted. They prove a library body is legal for a region. The price decides whether to use it.