What Weft is for, and where it loses
A compiler whose generated kernels are 5% slower than mine, whose pipelines are faster anyway, and whose losses are printed at the same size.
· 4 min · Ernest
Compilers & SIMD, Negative results, CorrectnessWeft started as a bet: that the place to win is not inside a kernel but in how a group of them is laid out, fused and scheduled for one particular machine, and that a compiler pricing those decisions against real measurements could beat what I'd pin by hand. The first honest result looked like the bet had lost. The kernels Weft generated came in about 5% slower than the NEON I'd written myself, and they still do. Then I measured the whole pipeline, and the version built from Weft's slightly-worse kernels finished in 92% of the hand-written one's time, because the compiler was better than I was at deciding what runs next to what. The win was orchestration. That's the result this project is about, and it took me a while to believe it.
So, in plain words. A computation graph is a model or a media pipeline written as operations and the data flowing between them. Weft takes a region of that graph, bigger than one operation and smaller than the whole thing, and decides, for the specific machine in front of it, how the region is laid out in memory, vectorised, fused and scheduled. Then it generates the code: .NET IL, C# source, Metal or CUDA. Every output is checked against a slow scalar reference and has to match it exactly. If a plan can't be verified it doesn't compile; there is no quiet fallback that emits scalar code and calls itself a backend. It's written in C# on .NET 10, which is not where people usually go for this.
The kernels lost. The pipeline won.
The first test was an H.264-like pipeline, a codec-shaped chain of stages. Hand-written here means NEON intrinsics I wrote myself, and the generated kernels are measured against them one at a time. On that scorecard the median generated kernel sits at about 1.05× the hand-written time. Five percent slower. On its own, a loss.
The whole pipeline, though: 9.191 µs per frame generated, against 9.983 µs for the hand-written sequential version, a ratio of 0.92×. The kernels didn't get better. The composition did. Weft binds stages together without copying between them and overlaps stages wherever the dependencies legally allow it, and no single kernel can see any of that. The project's own line for it: the win is orchestration, not suspicious benchmarks.
Eight percent is not a dramatic number. What makes it interesting is where it came from. I had been carefully optimising the part of the program that was already fine.
Nobody told the planner the answer
The second test was bigger: the 32-block vision tower of SAM 3, Meta's segmentation model, run as one graph.
| Before | After | Ratio | |
|---|---|---|---|
| Single-threaded | 1,394.1 s | 108.3 s | 12.9× |
| Eight workers | 77.0 s | 15.6 s | 4.93× (ceiling 5.11×) |
The 5.11× ceiling is what I reached by pinning the schedule by hand, and the gap to it belongs in the same sentence as the win. All 32 blocks came out bit-identical to the reference, and the source tree didn't change: nothing was added to make the model fit.
What I care about most is how the planner got there. It was fed only measured rows, how long each candidate actually took on this machine, never the answer, and it reached 96.5% of the ceiling I'd found by hand. Its evidence comes in tiers: an exact measurement, a nearby one, a model, a heuristic, and finally Unknown. Unknown is a real answer. Where the evidence doesn't exist, Weft says so rather than guess, and a compiler with a working representation of its own ignorance is a slightly unusual thing to own.
One clause on settings. Under the strict contract that forbids reordering float math, Weft is slower than ggml on this model; under the one a shipped model would actually use, faster. The tower numbers are the second.
Where it loses
Three places, printed at the same size as the wins.
On GPU, Weft's generated CUDA contraction runs at between 0.025× and 0.211× of cuBLAS: between 4.7× and 39× behind, same box, same problems, same runs. It is also bit-exact, against Weft's own CPU reference rather than against cuBLAS, so exact means the kernel is correct, not that it agrees with NVIDIA's. Correct and up to 39× slow is a code-generation problem, and I'd rather have the second number visible than bury it under the first.
On x64, the SAM 3 encode sits 3.754× behind ggml. That one is on weaker footing: the report gives the ratio but not the host, thread count or build flags for either side, so the direction is established and the size is indicative. I lost. I can't tell you exactly by how much.
And the one time I tried a Weft-generated kernel inside Kiln, my H.264 encoder, for quarter-pel luma interpolation, it measured 0.99× to 1.02×. A range that straddles 1.00 isn't a delta, it's a coin. The integration was declined and the branch never merged. A compiler project's most credible artefact may be the integration it tried and turned down.
So the bet isn't settled. Where Weft wins, it wins by scheduling. Two of the three losses are single kernels, one against cuBLAS and one against a kernel I'd already written by hand, which is bad for my ego and, annoyingly, consistent with the thesis: inside a kernel is exactly where I said the win wouldn't be. The third is a whole model on a box I didn't write down, which I need to go and measure properly.