Before you benchmark anything in .NET, run A against A
The same code, measured against itself, came out 19.99% apart. Until you know that number for your own harness, a 20% win is unreadable.
· 3 min · Ernest
MeasurementBefore I changed anything in the benchmark harness for Weft, a compiler of mine, I ran the dullest experiment I know. Same code, same configuration, split into two arms and timed as if one were the candidate and the other the baseline. The two arms came out 19.99% apart. Nothing was different between them. That number was the harness.
That's an A/A control. A normal benchmark is A/B: here's the old code, here's the new, which one is faster? An A/A control compares a thing with itself, so the true difference is exactly zero and whatever separation you get is your noise floor. It is the only way I know of to find out what your harness does when nothing is happening.
Mine, under .NET's default settings, did 19.99%. The setting that mattered was tiered compilation: the runtime ships a quick first compile and re-JITs hot code later, so the code you're timing can change under you partway through a run. With tiering off, the same control separated its two arms by 0.06% to 0.21%, and the coefficient of variation inside a process fell by 7 to 10 times.
DOTNET_TieredCompilation=0| Arm | A/A separation |
|---|---|
| default tiering | 19.99% |
| DOTNET_TieredCompilation=0 | 0.06% to 0.21% |
With tiering on, a real 20% effect and no effect at all look the same to this harness. That's a fact about the instrument first and the workload a distant second. Every sub-20% result I'd have reported before that control existed isn't wrong, necessarily. It's unreadable, which is worse, because wrong at least tells you something.
The second way the harness lied
Tiering wasn't the only thing my harness was doing to my numbers. One benchmark reported a sum-of-absolute-differences kernel at 15.16 ns on its smallest input. The real figure was 4.24 ns. The kernel was fine. The harness was allocating 24 bytes per call inside the timed region, from a LINQ FirstOrDefault in a name lookup, and charging that to the code under test. Harness overhead is close to fixed, so it hurts the shortest kernel most, and that is a 3.6× error on a benchmark with nothing wrong in it except the benchmark. The lookup is allocation-free now and every row reports zero allocations; there's a lab entry on it if you want the details.
Two rules fell out of all this. First, an effect under the noise doesn't get rounded up. A Kiln run once measured 2.290 ms against 2.235 ms with one flag flipped: 1.024×, a difference under twice the larger standard deviation, with the sample count not stated. I wrote it down as 1.024× and marked it not a result, because "about 3% faster" would have made a movement smaller than the noise look like a finding. Second, a delta that doesn't keep its sign between runs isn't a delta. It's an unstable reading, and rerunning it until it points the right way is how you manufacture one.
What to do with BenchmarkDotNet open
- Run A against A first. Same method, same inputs, two arms. Don't change anything until you've seen the separation; that number is your floor.
- If the floor is wide, run the control again with DOTNET_TieredCompilation=0 and see what it does to the separation. Mine went from 19.99% to under a quarter of a percent.
- Make the benchmark report allocations, and expect zero inside the timed region. If it isn't zero, the harness is in your measurement.
- Anything under the floor gets reported as not a result, at the precision where it's visible, and doesn't get retried into existence.
An A/A control is boring on purpose. It can't tell you your change worked. It tells you whether your harness could have told you, and for a while the honest answer in my case was no.