Lab / / benchmark
4.93× against a 5.11× ceiling
The SAM3 ViT tower went from 1,394.1 s to 108.3 s single-threaded and 77.0 s to 15.6 s at W=8, with 16,671,744 elements bit-identical and `src/` untouched.
SIMD and intrinsics, Benchmarking, CorrectnessWeft- Machine
- UNKNOWN — .NET 10 stated, machine not
- Commit
- UNKNOWN
- Single-thread, before → after
- 1,394.1 s → 108.3 s
- W=8, before → after
- 77.0 s → 15.6 s
- Single-thread ratio
- 12.9×
- W=8 ratio against the hand-pinned ceiling
- 4.93× against 5.11×
- Blocks bit-identical
- 32 of 32
- Elements bit-identical
- 16,671,744
- Detector graph
- 831 checkpoint tensors, 819.5 MiB
The SAM3 32-block ViT tower went from 1,394.1 s to 108.3 s single-threaded, a factor of 12.9. At W=8 it went from 77.0 s to 15.6 s, a factor of 4.93.
The 4.93× is quoted against a 5.11× hand-pinned ceiling rather than against the starting point. The gap to the best achievable configuration sits in the same sentence as the win, which is the honest way to state it and the way I want it copied.
SAM3 32-block ViT tower
| Measurement | before | after | Delta | Ratio |
|---|---|---|---|---|
| single-thread (s) | 1,394.1 | 108.3 | faster | 12.9×quality unknown |
| W=8 (s) | 77.0 | 15.6 | faster | 4.93× (against a 5.11× ceiling)quality unknown |
Why `src/` untouched matters
32 of 32 blocks are bit-identical. Over the full detector graph, 831 checkpoint tensors are held by reference and 16,671,744 elements come out bit-identical, with `src/` untouched.
That last phrase carries the claim. Sixteen million elements identical while the source tree does not change means the improvement is in generated code rather than in a hand-edit. It is the cleanest place Weft gets to make that claim.