Lab / AUG 25 2026 / experiment
The metric that moved the wrong way
Jitter was backwards: the clips that sound fried have more stable pitch, 10.3% F0 deviation against 13.5% and 14.7%. Period-to-period inconsistency separated the populations and still did not fix the percept.
Audio, Negative results, Measurement methodologyModalis- Machine
- UNKNOWN — "dev machine", unnamed
- Commit
- `ddc1eb9`, 2026-08-25
- F0 deviation, fried clips
- 10.3%
- F0 deviation, good clips
- 13.5% and 14.7%
- Cohen's d, period-to-period inconsistency
- 2.08
- AUC, period-to-period inconsistency
- 0.983
- Churn, exp_018
- 6.58 → 3.62 (−45%)
- Consistency, exp_020 quietest-bed quartile (n=14)
- 0.845 vs 0.541 busiest; r = −0.394
Jitter was the metric I expected to find the problem, and it moved the wrong way. The clips that sound fried measured 10.3% F0 deviation. The good clips measured 13.5% and 14.7%.
The clips that sound worse have more stable pitch. Two other hypotheses were refuted the same way — low F0, and timing irregularity — leaving per-pulse shape, which was also refuted. Three of the four measures refuted in Gate 2 were refuted by later measurement rather than by argument.
What did separate them
Period-to-period inconsistency separated the two populations almost perfectly: Cohen's d 2.08, AUC 0.983. An engine built to repair it did not produce audible improvement. A statistic that discriminates is not a model that generates.
| Technique | Objective result | Verdict |
|---|---|---|
| WORLD resynthesis | jitter −44%, HNR 2.8→7.9 dB | "still sounds fried" |
| LPC excitation | jitter −36% | no improvement |
| PSOLA | jitter −76% | "perception only marginally better" |
| Pulse consistency | consistency 0.57→0.87 | "muffled as shit" |
Gate 2's answer is a single word — *"Answer: No."* — and the techniques above are why. Every one of them moved its target metric. None of them fixed the thing the metrics were standing in for.