Writing

Weft now writes its own machine code

Until September, RyuJIT compiled every kernel Weft produced. Now Weft can emit x64 and arm64 directly, checked bit for bit against the kernel it replaces. Why I did it, what it took, and what it bought in week one.

28 Sept 2026 · 7 min · Ernest

Compilers & SIMD, Correctness

Weft generates its kernels as .NET IL and hands them to RyuJIT, the runtime's JIT, which does instruction selection and register allocation and mostly does them well. That was the design, and I'd still defend it: on a fused elementwise chain the generated code matched my hand-written NEON at 1.00×. But over the summer three things kept turning up in listings that I couldn't fix from the IL side.

The first is instructions .NET can't reach. Widening half-precision to float is one instruction on both ISAs, VCVTPH2PS on x64 and FCVTL on arm64, and .NET 10 has no route to either: no F16c class, no Half overload on Widen, no Half-typed member anywhere in the intrinsics namespace. I checked by reflection and then by disassembly, because I didn't believe it. Weft's shipped widening is eighteen vector instructions and a store. Standing a reachable one-op conversion in for the missing one bounded what it's worth at 1.08× on a 28-block Q8_0 kernel, which is not a lot, except that every quantised weight in a language model goes through it. arm64's int8 matrix instruction, SMMLA, has no intrinsic either.

The second is register allocation I couldn't steer. RyuJIT folds a static readonly Vector256 constant into the code only if its class's constructor has already run when the method is jitted. With measured selection on, the fused attention body happened to be jitted before the exp provider's constructor, so its exp loop carried a class-init check and twelve constants as 64-bit addresses, ten of them spilled and reloaded from the frame on every eight-lane group. The fix was to spell constants as const scalars broadcast at the point of use, plus a test that refuses any static readonly SIMD field in a shipped assembly. It failed on master naming 23 of them. The old conv tile spilled ten of its twelve accumulators every k-step for an unrelated reason. All of this was fixable in C#, and all of it was found by reading machine code I didn't write, after the fact.

The third is pricing. Weft's planner chooses between kernels by cost, and the cost should be of the code that will actually run. I built a spike that captures RyuJIT's emitted bytes in-process and prices them with a port-level model, and it landed about three times closer to the measured wall than the roofline floor. Useful, but you only know the bytes after the JIT has produced them. If Weft emits the bytes, it knows them before anything runs.

Twins, not a backend

So the native engine is deliberately not a new backend. It's a seam, an emitter interface, and a set of twins. A twin takes the managed kernel's own offer (its contract, its price, its packed operands, its member views for threaded splits) and replaces only the body with machine code. Before the twin is offered to the planner it's validated bit for bit against the arm it mirrors. A mismatch is a refusal with a name, and in the developer-override mode a mismatch stops the compile. The twin follows the managed listing's instruction and operand order, so the arithmetic rounds identically. On x64 the bytes come from the Iced assembler. On arm64 I wrote the A64 encoder, and every instruction form it knows is checked against clang's rendering of the same text. Each native kernel is a leaf: one pointer to an argument struct, called through an unmanaged function pointer with the GC transition suppressed. It never throws and never allocates.

Executable memory is W^X. Map a page read-write, copy the code in, flip it to read-execute, and the region is never writable again; whether the host will allow that at all is probed before any code is placed. Then macOS. Apple silicon wants MAP_JIT with a per-thread write toggle, and toggling it from managed code crashed the CI test host on the M4, because the managed caller is itself running out of the runtime's MAP_JIT heap. The first version refused macOS under the signed dotnet host on the grounds that it needed an entitlement the host didn't carry. It didn't. The actual fix is a small C shim, built from source at build time and never committed as a binary, that opens the toggle, writes, and closes it inside one native call, and does the arm64 instruction-cache maintenance on Linux as well. While building it I found the old probe hadn't even been a safe gate on the M4: a plain anonymous page flipped to read-execute reports success from mprotect, and the process is killed on its first instruction. Windows got VirtualAlloc and the Win64 calling convention, and the first bodies that assumed System V crashed the Windows test host before anything got round to refusing them by name.

On 27 September native emission went on by default. Every family of twins registers as an ordinary candidate: validated, then priced against the managed arm by the cost model, with the IL path still in place behind it. A twin is priced from its own bytes: decode them, find the loops, multiply each instruction by its trip counts at the kernel's geometry, and run the port model where the host has a table. An exact tie keeps the shipped arm, and that rule was measured in rather than designed in. With the tie decided after selection, a same-priced twin sitting in the roofline's field turned 144 Q8_0 matrix-vector kernels on the M4 from the int8 NEON arm into Apple's sgemm: 9.04 ms per token instead of 8.77, and a different logits digest. Verdicts are cached per host so the validation runs once. WEFT_NATIVE_EMIT=off is the kill switch and gives a byte-identical compile.

What it bought in week one

On the kernels that already had good managed bodies, about what the design predicted, which is to say not much.

WhatMachineManagedNativeNative ÷ managed
Qwen decode, Q8_0 GEMV twinNucBox, Zen 418.18 ms/token17.84 ms/token0.981 (A/A 1.004)
Qwen decode, SMMLA twin of the int8 SDOT armM410.07 ms/token10.58 ms/token1.051 (A/A 1.001)
Strict softmax-exp, correctly-rounded NEON twin vs the generated bodyM44,564 µs2,086 µs0.457
Native twin against the managed arm it replaces, same machine, interleaved processes, an A/A pair alongside. Qwen is the 0.5B model's decode at one thread; the softmax-exp is a [201 × 5184] masked row under Strict. The NucBox is a Ryzen 9 8945HS (Zen 4); the M4 is a Mac mini. Late September.

The x64 Q8 twin is 2% faster at the decode, on 1,014 of 1,014 Q8_0 kernel rows, with the same 64 tokens and the same logits digest. RyuJIT's losses on that arm were address arithmetic and one spilled bound, not the arithmetic, and 2% is what that's worth. The SMMLA twin on the M4 lost 5%: at one row it does the same useful work per instruction as the SDOT arm and pays four ZIPs per column pair plus a duplicated activation. The two-row interleave that would remove the ZIPs isn't built. A twin that's bit-exact and slower stays in the catalogue, because it's priced, so it isn't selected.

The wins are where RyuJIT had nothing to offer. Strict's exp is now a correctly-rounded one on every platform (glibc's own expf misses correct rounding on 170,648 of the 2³² inputs and Apple's on two million, and the whole-domain output hash is identical on both boxes), and as a NEON twin it runs 2.2× faster than the generated body. Fused online-softmax attention exists on arm64 for the first time, because the x64 one was a managed body the arm64 side never had; per head at SAM 3's shape, one thread on the M4, 41.9 ms native against 239.4 ms for the six kernels it replaces, and all 632 attention chains in the SAM 3 plan now render native on both ISAs. The Q8 twins widen their fp16 deltas with F16C, the instruction from the second paragraph. The matmul and conv twins read 20 to 30% faster than their arms on x64, but that box was running someone else's CI at the time, so those are directions, not numbers.

Twins exist today for the hand-written arms: the seven x64 Q8_0 matrix-vector kernels, the seven F32 matmul arms on both ISAs, the direct convolutions, attention, and the elementwise exp and erf. The generated kernels, the ones the IL path builds from the plan's own loop nests, still go through RyuJIT; their native lowering is in review on both ISAs, bit-identical to the IL path. The other open lane prices a twin against its arm in cycles from its own bytes and runs a cached micro-race when the model can't separate them, which is the same idea the composition engine uses for regions. The IL path isn't going anywhere. It's what every twin is checked against.