Lab / / failure

Generated CUDA at 0.025× of cuBLAS

Weft's generated CUDA contraction runs between 4.73× and 39.38× behind cuBLAS, and is bit-exact against Weft's own CPU oracle.

GPU, Negative results, CorrectnessWeft
Machine
8× RTX 5090, sm_120; CUDA 13.1, driver 590.48.01
Commit
UNKNOWN
Weft CUDA contraction vs cuBLAS
0.025× – 0.211×
Expressed as a deficit
4.73× to 39.38× behind
Correctness
bit-exact against Weft's own CPU oracle

Weft's generated CUDA contraction runs at 0.025× to 0.211× of cuBLAS. That is between 4.73× and 39.38× behind.

It is also bit-exact against Weft's own CPU oracle.

The pairing is the finding

The generated code is correct and it is up to 39× slow. I am reporting those two facts in the same paragraph because they are separate axes, and this is the clearest case in the work where reporting only the flattering one would have been easy. Codegen that produces the right answer and runs 39× behind is a codegen problem, not a numerics problem, and it needs the second number to be visible for anyone to work on it.

MeasurementcuBLASWeft generated CUDA contractionDeltaRatio
contraction throughput, worst case1.00× (reference)0.025×39.38× behind0.025×quality unknown
contraction throughput, best case1.00× (reference)0.211×4.73× behind0.211×quality unknown
8× RTX 5090, sm_120; CUDA 13.1, driver 590.48.01. Date and commit UNKNOWN.Same box, same problems, same runs — MATCHED. Direction: candidate ÷ baseline, where higher is better; both rows sit below 1.00, so both are losses. Correctness is a separate axis: the output is bit-exact against Weft's own CPU oracle.Not every row is a matched comparison — read the quality on each rationo source recordedUnverified · LOW

The range is wide because the ratio depends on the shape. Both ends of it are losses, so I am not quoting the top of the range as the result.