Projects / experimental
SAM3 kernels
CUDA and SIMD kernel work on a real detector, where the same measurement runs in opposite directions under two contracts.
- Status
- experimental
- Languages
- C#, CUDA, C++
- Current release
- —
- Latest
- —
Strict contract 3.23–3.31x slower
Inference contract 2.9751x faster
Same detector, same machine, same ggml anchor. The posture is the missing denominator — see BENCHMARK_LEDGER.md weft-sam3-strict.
How it works
CUDA and SIMD kernels for a segment-anything-class detector, measured against a ggml anchor under two floating-point contracts.
SAM3 is work on the kernels of a segment-anything-class detector, benchmarked against a ggml anchor. It produced the survey's cleanest demonstration that a benchmark ratio without its contract is not a measurement: the same detector, on the same machine, against the same anchor, is 3.23–3.31x slower under the Strict contract and 2.9751x faster under Inference.
- The same detector, machine and anchor: 3.23–3.31x slower under Strict, 2.9751x faster under Inference.
- Neither number is wrong, and neither is 'the' number — the contract is the missing denominator.
- A CUDA contraction-loss result recorded as a loss.