Kiln: an H.264 encoder in C#, and what the spec didn't tell me
Writing a video encoder with no native code underneath turned out to be possible. Knowing when it was wrong was the harder part.
· Updated 1 Oct 2026 · 4 min · Ernest
Media & codecs, CorrectnessProxeno, my game-streaming server for emulators, needed to encode H.264 inside a .NET process, and every option I found was one of two things: a native codec to cross-compile, ship and keep patched for every platform, or a GPL or LGPL one I didn't want to link. There is no clean-licensed, dependency-free, low-latency H.264 encoder for .NET. So I wrote one. Kiln is pure managed C#, no native code underneath, implemented clean-room against the ITU-T H.264 spec. Feed it raw frames and it hands back a standards-compliant baseline bitstream.
Baseline profile only, on purpose. No B-frames, no CABAC, no 8×8 transform, no interlace; 4:2:0, 8-bit. That is the subset a low-latency stream actually uses, and it is where the README stops: Kiln is not an x264 competitor and isn't trying to be. The goal was a stream I could produce from managed code without shipping anyone else's binary.
Wrong slowly
Here's the part I didn't expect. Speed wasn't the hard problem. The hard problem was being wrong in a way nothing complained about.
Deblocking is the filter that smooths block edges after reconstruction, and encoder and decoder have to apply it identically, because the encoder predicts the next frame from the filtered picture it believes the decoder has. I had two bugs in it. The luma filter ignored the spec's rule to average QP across a macroblock edge, and my boundary-strength derivation had its precedence backwards: coded coefficients are supposed to win over motion-vector differences, and mine let the motion vectors win.
Neither bug produced an invalid stream. Every decoder played it. My reconstruction just drifted a little further from every conformant decoder's with each frame, because the encoder was predicting from a picture nobody else had. Nothing failed. I only found out because Kiln is built against golden reference frames generated by ffmpeg, and every stream it produces is decoded by ffmpeg and by Apple's VideoToolbox, two decoders I didn't write, and compared against what they reconstruct. A test I'd written to match my own reading of the spec would have passed, because my reading was the bug.
After the fix, reconstruction is byte-exact against both. Some streams now produce different bytes: the ones using per-macroblock QP, and constant-QP streams at QPs where the spec distinguishes two boundary strengths, such as 31, 32, 35 and 36. That's documented as a behavioural change, with the spec sections cited, so anyone can check the fix against the standard rather than against me. Default-option streams and constant QP 23, 28, 33 and 34 were never affected, which is a smaller blast radius than "the deblocking filter was wrong" sounds like.
Four slices buy 2.2×
A slice is a band of the frame that can be encoded independently, so slices are how an encoder uses more than one core, and how a lost packet stays confined to part of a picture. The obvious expectation is that four slices make a frame about four times faster. These numbers are from an Apple M5 Max under .NET 10, QP 28, steady P-frames over synthetic content, with the competing configurations interleaved in one process so scheduling drift hits every arm equally:
| Slices | Median per frame |
|---|---|
| 1 | 24.4 ms |
| 2 | 16.9 ms |
| 4 | 10.9 ms |
| 8 | 10.8 ms |
Four slices buy about 2.2×, not 4×, and eight buy nothing beyond four. Per-slice motion cost doesn't balance perfectly, and every frame still pays serial work that no amount of slicing removes. Slices also cost bits, because a slice boundary resets motion-vector and skip prediction: +5% to +26% bitrate at QP 28 going from one slice to four, depending on content. Pick a slice count for latency and loss confinement, then budget for both. Linear returns are not on offer.
The number I'd want you to read first
Medians are what people size a deployment from, and here they are the wrong number. On a divergent-motion stress generator (motion going a different way in every part of the frame), the default preset measures 205 ms per frame at 1080p single-slice, against 27 ms on coherent content in the same run. A sevenfold swing, decided by the content, with no bound. The default is the only preset that leaves per-macroblock motion-search effort uncapped, so it is the only one with no worst case. A deployment sized from the 10.9 ms four-slice median meets a 70 ms frame on that content and drops to 14 fps.
The fix is in the README: keep the default quality and set an explicit effort cap. On coherent content that costs about −0.01 dB. On the hostile content it pulls the four-slice frame in to roughly 35–40 ms. The cap counts algorithmic work, never wall clock, so the bitstream stays deterministic. That is a determinism guarantee, not a performance one.
Kiln is public on NuGet as Proxeno.Kiln, Apache-2.0, pre-release 0.x; the APIs will change. There are 2,292 tests, and I'd still rather you read the worst-case row than the test count. The spec told me how to filter a block edge. It didn't tell me that getting it slightly wrong would look exactly like getting it right, for as long as I was the only one checking.