Lab

Notebook

Shorter than an essay. Dated. A failure is allowed to stay a failure.

  1. SEP 17 2026

    benchmark

    A result that survives a contaminated box

    • Machine RTX 5090, Vast.ai instance 51353366, driver 580.173.02
    • Commit `gabilan/sam3-ggml` `ewi1963/f16-vectorized-converts` (`4f023a25`, `14a301d5`, `429236ef`)
    • frame, base → f16c 89.43 → 82.16 ms, −7.27 ms (−8.1%), 3/3
    • encode, base → f16c 53.57 → 48.22 ms, −5.35 ms (−10.0%), 3/3
    • prop, base → f16c 35.85 → 34.11 ms, −1.74 ms
    • other_cpus across rounds 3.59 → 10.20
    • base frame spread across rounds 89.43 → 89.27 → 89.56 (0.29 ms)
    • f16c frame spread across rounds 82.16 → 82.67 (0.57 ms)
    • Bit-identical records, base vs f16c 80/80, worst-cell IoU 1.000000, max|Δlogit| 0

    Every run read gate=violated while interference varied threefold. The arms did not drift with it: frame 89.43 → 82.16 ms, encode 53.57 → 48.22 ms, 3/3 on the sign, 80/80 bit-identical.

  2. SEP 17 2026

    benchmark

    The second instrument

    • Machine RTX 5090, Vast.ai instance 51353366
    • Commit kernels on `gabilan/sam3-ggml` branches; census harness has no commit
    • Total GPU kernel time / encode, base → f16b → f16c 45.78 → 40.66 → 40.22 ms (Δ −5.56 ms)
    • kernels that moved 5
    • kernels flat within noise 26
    • im2col, base → f16c 2.32 → 1.87 ms/encode (−19%)
    • Iteration 3 frame delta vs within-arm spread +0.48 / +0.77 / +0.41 ms vs 0.44 / 0.57 ms

    The wall clock could not resolve iteration 3. The kernel census could: total GPU kernel time per encode 45.78 → 40.22 ms, with im2col down 19%.

  3. SEP 17 2026

    note

    Two arms drifted +11 ms together while every gate read ok

    • Machine RTX 5090, Vast.ai instance 51353366
    • Commit `gabilan/sam3-ggml` `ewi1963/f16-vectorized-converts@429236ef` (kernels); the harness has no commit
    • Co-drift in the previous campaign +11 ms, both arms together, every gate reading OK
    • Interference variation in this campaign other_cpus 3.59 → 10.20

    A previous campaign's arms co-drifted by 11 ms with every environment gate reporting OK. The response was not a better gate but an interleaved protocol with the per-round series printed.

  4. SEP 10 2026

    experiment

    The A/A control that said 19.99%

    • Machine UNKNOWN — not stated in the source
    • Commit UNKNOWN
    • A/A separation, default tiering 19.99%
    • A/A separation, DOTNET_TieredCompilation=0 0.06% – 0.21%
    • In-process coefficient of variation falls 7–10×

    Measuring a configuration against itself turned up a 19.99% separation under the default .NET runtime. With tiering off it fell to 0.06–0.21%.

  5. AUG 25 2026

    failure

    None was caught by the numbers looking wrong

    • Machine UNKNOWN — "dev machine", unnamed
    • Commit `ddc1eb9`, 2026-08-25
    • Measurement errors reported 6
    • Caught by listening alone 4 of 6

    Six measurement errors in one gate. Every one returned plausible numbers; four of the six were caught by listening alone.

  6. AUG 25 2026

    experiment

    The metric that moved the wrong way

    • Machine UNKNOWN — "dev machine", unnamed
    • Commit `ddc1eb9`, 2026-08-25
    • F0 deviation, fried clips 10.3%
    • F0 deviation, good clips 13.5% and 14.7%
    • Cohen's d, period-to-period inconsistency 2.08
    • AUC, period-to-period inconsistency 0.983
    • Churn, exp_018 6.58 → 3.62 (−45%)
    • Consistency, exp_020 quietest-bed quartile (n=14) 0.845 vs 0.541 busiest; r = −0.394

    Jitter was backwards: the clips that sound fried have more stable pitch, 10.3% F0 deviation against 13.5% and 14.7%. Period-to-period inconsistency separated the populations and still did not fix the percept.

  7. AUG 21 2026

    benchmark

    1.024×, and not rounded up

    • Machine arm64, `RELEASE_ARM64_T6050`, Darwin 25.6.0
    • Commit UNKNOWN — the artifact records no commit
    • EnableIntraInPFallback=False 2,290,059 ns (2.290 ms), stddev 28,141 ns
    • EnableIntraInPFallback=True 2,235,369 ns (2.235 ms), stddev 23,746 ns
    • Ratio 1.024×
    • Difference 54.7 µs

    The intra-in-P fallback measured 2.290 → 2.235 ms, a 1.024× movement. The difference is under twice the larger standard deviation, so it is recorded and marked as not a claim.

  8. experiment

    A control that proves the faults landed

    • Machine UNKNOWN — the integration tests do not state one
    • Commit `cacb70c`, 2026-08-25
    • Stream 2,098 packets
    • Permanent holes, retransmission off 19 / 113 / 314 at 1% / 5% / 15% loss
    • Repairs, retransmission on every packet the receiver could detect as missing

    "Every recoverable packet came back" is a claim a broken injector could also make. Running the same seeds with retransmission off left 19 / 113 / 314 permanent holes.

  9. benchmark

    A 16×16 window at 9.92 ns against 4.85 ns by hand

    • Machine Apple M5 Max, arm64
    • Commit UNKNOWN
    • 16×16, hand strided idiom 4.85 ns
    • 16×16, weft provider 9.92 ns
    • 16×16, weft generic scalar nest 94.6 ns
    • 8×8, hand strided idiom 1.59 ns
    • 8×8, weft provider 6.43 ns
    • 8×8, weft generic scalar nest 24.2 ns

    The shape motion search actually computes: 16×16 at 4.85 ns by hand against 9.92 ns from the provider, with the 8×8 comparison carrying a kernel-shape difference.

  10. failure

    Generated CUDA at 0.025× of cuBLAS

    • Machine 8× RTX 5090, sm_120; CUDA 13.1, driver 590.48.01
    • Commit UNKNOWN
    • Weft CUDA contraction vs cuBLAS 0.025× – 0.211×
    • Expressed as a deficit 4.73× to 39.38× behind
    • Correctness bit-exact against Weft's own CPU oracle

    Weft's generated CUDA contraction runs between 4.73× and 39.38× behind cuBLAS, and is bit-exact against Weft's own CPU oracle.

  11. observation

    Two frames in three seconds, and a mechanism for the rest

    • Machine UNKNOWN — "the dev machine"
    • Commit UNKNOWN
    • Frames drawn in a 3 s idle run 2
    • First frame after process start 187–450 ms
    • Derived claim about one frame per minute

    The project states 2 frames drawn in a 3 s idle run and derives one frame per minute from it. The mechanism is committed; the number is not reproducible from the repo.

  12. benchmark

    4.93× against a 5.11× ceiling

    • Machine UNKNOWN — .NET 10 stated, machine not
    • Commit UNKNOWN
    • Single-thread, before → after 1,394.1 s → 108.3 s
    • W=8, before → after 77.0 s → 15.6 s
    • Single-thread ratio 12.9×
    • W=8 ratio against the hand-pinned ceiling 4.93× against 5.11×
    • Blocks bit-identical 32 of 32
    • Elements bit-identical 16,671,744
    • Detector graph 831 checkpoint tensors, 819.5 MiB

    The SAM3 ViT tower went from 1,394.1 s to 108.3 s single-threaded and 77.0 s to 15.6 s at W=8, with 16,671,744 elements bit-identical and `src/` untouched.

  13. experiment

    A flat result, reported flat

    • Machine Apple M5 Max, arm64
    • Commit UNKNOWN
    • 1 accumulator chain 356.3 ns
    • 2 accumulator chains 356.2 ns
    • 4 accumulator chains 356.6 ns

    One, two and four accumulator chains measured 356.3 / 356.2 / 356.6 ns. The dependency-chain argument does not apply to this kernel shape.

  14. observation

    Rows nobody retired, stale by 5.76×

    • Machine Apple M5 Max, arm64
    • Commit UNKNOWN
    • Superseded rows 13,071.574 ms
    • Current row 2,268.278 ms
    • Staleness of the superseded rows 5.76×
    • Pinned ggml 1T anchor 6,748.36 ms

    Every SAM3 1T row below the warning in the scorecard predates the 2026-09-02 single-thread push and is stale by 5.76×. They were kept, not edited.

  15. failure

    Struck as UNMATCHED, then falsified at 1.24× slower

    • Machine UNKNOWN — not stated
    • Commit UNKNOWN
    • Original claim 1.29× faster than Meta
    • Matched-protocol result 1.24× slower

    A "1.29× faster than Meta" claim was produced from an unmatched comparison, quoted as a prize, struck, and then inverted to 1.24× slower on a matched protocol.

  16. failure

    A bake-off invalid by construction: N_prop 80 vs 16

    • Machine UNKNOWN — not stated
    • Commit UNKNOWN
    • N_prop, Weft/ggml (C++) side 80
    • N_prop, Meta reference side 16

    The two sides of the SAM3 bake-off did different work per row: N_prop 80 on one, 16 on the other. No ratio from it is published.

  17. benchmark

    Five minutes in, less heap than at the start

    • Machine UNKNOWN — not stated
    • Commit `cacb70c`, 2026-08-25
    • Duration and rate 5 minutes at 30 fps
    • Link 5% lossy, reordering
    • Packets 31,815
    • Repaired 1,564
    • Holes 0
    • History misses 0
    • Managed heap at end 0.9% below where it started

    A five-minute soak at 30 fps over a 5% lossy, reordering link moved 31,815 packets with zero holes and ended with managed heap 0.9% below where it started.

  18. observation

    The screenshot tool reports success while the panel stays blank

    • Machine UNKNOWN — Pi vc4 is named as the failing target
    • Commit UNKNOWN
    • Number of frames involved none — this is a mechanism, not a measurement

    `SDL_WINDOW_FULLSCREEN_DESKTOP` makes every `eglSwapBuffers` fail on Pi vc4. GLES still renders, so `--screenshot` produces a real image of something the panel never receives.

  19. failure

    3.754× behind ggml, on weaker footing

    • Machine UNKNOWN — not tied to a host
    • Commit UNKNOWN
    • SAM3 x64 encode vs ggml 3.754× behind
    • Qwen vs llama.cpp, end to end 6.18× behind

    SAM3 x64 encode measured 3.754× behind ggml and Qwen 6.18× behind llama.cpp. The report gives the ratios but not the hosts, thread counts, or build flags.

  20. benchmark

    433 frames versus 51, same seed

    • Machine UNKNOWN — the client was headless Chrome 151
    • Commit `cacb70c`, 2026-08-25
    • Frames decoded, with RFC 4588 RTX 433
    • Rate, with RTX 29 fps
    • Freezes, with RTX 0
    • Packets accepted back, with RTX 79 of 80 destroyed
    • Frames decoded, without 51
    • Rate, without 15 fps
    • Freezes, without 5

    The same fifteen seconds over a 5% lossy link, twice, same seed: 433 frames / 29 fps / 0 freezes with RFC 4588 RTX against 51 frames / 15 fps / 5 freezes without.

  21. failure

    A result that did not keep its sign

    • Machine UNKNOWN — not stated
    • Branch `weft-pilot/motion-sad` (left unmerged)
    • Reported speedup 0.99× – 1.02×
    • Circulating claim that does not reproduce ~50% faster with Weft

    qpel luma interpolation through a Weft-generated kernel measured 0.99×–1.02×. The integration was declined and the work stayed on its branch.

  22. failure

    24 bytes inside the timed region

    • Machine Apple M5 Max, arm64
    • Commit UNKNOWN
    • Corrected, N=256 4.24 ns
    • Earlier invalid run, N=256 15.16 ns
    • Distortion 3.6×
    • Allocated per operation, inside the timed region 24 bytes

    An earlier SAD run reported the idiom at 15.16 ns instead of 4.24 ns. The cause was 24 bytes of allocation per operation inside the measured region.