Lab / SEP 17 2026 / benchmark
A result that survives a contaminated box
Every run read gate=violated while interference varied threefold. The arms did not drift with it: frame 89.43 → 82.16 ms, encode 53.57 → 48.22 ms, 3/3 on the sign, 80/80 bit-identical.
Contaminated runs, GPU, Measurement methodology, CorrectnessSAM3- Machine
- RTX 5090, Vast.ai instance 51353366, driver 580.173.02
- Commit
- `gabilan/sam3-ggml` `ewi1963/f16-vectorized-converts` (`4f023a25`, `14a301d5`, `429236ef`)
- frame, base → f16c
- 89.43 → 82.16 ms, −7.27 ms (−8.1%), 3/3
- encode, base → f16c
- 53.57 → 48.22 ms, −5.35 ms (−10.0%), 3/3
- prop, base → f16c
- 35.85 → 34.11 ms, −1.74 ms
- other_cpus across rounds
- 3.59 → 10.20
- base frame spread across rounds
- 89.43 → 89.27 → 89.56 (0.29 ms)
- f16c frame spread across rounds
- 82.16 → 82.67 (0.57 ms)
- Bit-identical records, base vs f16c
- 80/80, worst-cell IoU 1.000000, max|Δlogit| 0
Every run in the A/B reads `gate=violated`. Not one or two of them. `other_cpus` swings from 3.59 to 10.20, a threefold variation in interference.
The arms do not drift with it. Across three rounds the base arm reads 89.43 → 89.27 → 89.56 ms, a spread of 0.29 ms. f16b reads 82.64 → 83.08 and f16c 82.16 → 82.67, a spread of 0.57 ms. The measurement held steady while the environment varied threefold.
The result
Frame time moves from 89.43 to 82.16 ms, down 7.27 ms or 8.1%, on 3/3 paired rounds. Encode moves from 53.57 to 48.22 ms, down 5.35 ms or 10.0%, also 3/3.
The direction is worth quoting. Every paired row is labelled `paired frame (base-f16c, + = f16c faster)`, so a positive paired delta and a negative delta-in-milliseconds both mean faster, and f16c is the candidate in every row. The raw base → f16c series is +7.27 +7.17 +6.89. Assume the other convention and every figure reads backwards.
| Metric | base | f16b (iter 1+2) | f16c (iter 1+2+3) | base → f16c | sign |
|---|---|---|---|---|---|
| frame med (ms) | 89.43 | 82.87 | 82.16 | −7.27 ms (−8.1%) | 3/3 |
| encode med (ms) | 53.57 | 48.76 | 48.22 | −5.35 ms (−10.0%) | 3/3 |
| prop med (ms) | 35.85 | 34.11 | 34.11 | −1.74 ms | — |
**Environment.** RTX 5090, Vast.ai instance 51353366; `cgroup cpu.max: 2304000 100000` (2.304 cores); `threads=4`, `warmup=2 timed=5`, `n=400` frames, `MODE=1` (serial); input `sam3.1-multiplex-f16.ggml`, `N_prop=80`, `K=2`.
**Methodology.** `run_ab.sh`: three-arm interleaved A/B, 3 paired rounds, arms rotating position inside each round, a separate source tree and binary per arm, `BUILD_RC=0` on all four builds. Same machine, model, clip and K — MATCHED.
Correctness
The same-arm control is scored first, because cross-arm bit-identity is the one claim a mis-specified dump path can manufacture for free. base against base2 — the same binary, a second dump — reads 80/80 bit-identical with `max|Δlogit| = 0`, so the comparator's floor on this box is exactly zero. base against f16c is also 80/80, worst-cell IoU 1.000000.