Lab / AUG 21 2026 / benchmark

1.024×, and not rounded up

The intra-in-P fallback measured 2.290 → 2.235 ms, a 1.024× movement. The difference is under twice the larger standard deviation, so it is recorded and marked as not a claim.

Benchmarking, Negative results, Measurement methodologyKiln
Machine
arm64, `RELEASE_ARM64_T6050`, Darwin 25.6.0
Commit
UNKNOWN — the artifact records no commit
EnableIntraInPFallback=False
2,290,059 ns (2.290 ms), stddev 28,141 ns
EnableIntraInPFallback=True
2,235,369 ns (2.235 ms), stddev 23,746 ns
Ratio
1.024×
Difference
54.7 µs

The intra-in-P fallback measured 2,290,059 ns with it off and 2,235,369 ns with it on. That is 2.290 ms against 2.235 ms — a 1.024× movement, or 2.4% faster with the fallback enabled.

Armmean (ms)stddev (ns)
EnableIntraInPFallback=False2.29028,141
EnableIntraInPFallback=True2.23523,746
Encode_primmed_P, whole encode arm64, `RELEASE_ARM64_T6050`, Darwin 25.6.0 (macOS 26); the committed BenchmarkDotNet artifact `kiln/perf/h264-simd-perf-baseline-latest.json`, captured 2026-08-21T17:42:40Z. Commit: UNKNOWN. MATCHED — same run, one boolean differs. Direction: lower (faster) is better. `n` is not stated in the artifact, and the difference is under 2× the larger standard deviation.

I am reporting it as 1.024× rather than 1.03× or "about 3%" because the second decimal is the only place the effect is visible. Rounding it up would make a movement smaller than the noise look like a result.

I would not publish it as a result on its own. The difference of 54.7 µs is under twice the larger standard deviation of 28,141 ns, and `n` is not stated in the artifact. It is recorded because it is in the committed baseline, and marked as not a claim for exactly that reason.

The same artifact contains what a result looks like when the effect is bigger than the noise: `Satd4x4_Once` at 15.988 ns with intrinsics off against 11.182 ns with them on, a 1.43× speedup with a standard deviation of 0.038 ns on the fast arm.