Timeouts are not money limits
A per-hour cap on rented GPUs bounds the rate, not the bill. What I changed the day I noticed, and why I deleted the whole thing five days later.
· 4 min · Ernest
Correctness, MeasurementI have a small gateway in front of a locally hosted LLM. It checks who is calling, counts the request against their daily allowance, queues when the model is busy, and forwards. Boring on purpose. For a few days it also carried a control loop that could rent GPUs by the hour from vast.ai when the queue got long. The loop had a setting called MaxFleetPricePerHour, and on the day I moved it into its own service I read that name properly for the first time.
A per-hour price cap bounds the rate. It does not bound the bill. The realistic failure isn't an expensive instance; it's a cheap one that provisions, fails its readiness check, gets destroyed, gets re-rented, and does that in a loop. A circuit breaker would open eventually. What it cost between the first failure and the breaker opening was bounded by nothing at all. Every guard in that loop had a time dimension. None had a dollar dimension. Timeouts are not money limits.
Commit the worst case before you spend it
The fix is less clever than the bug. When a lease is created, the service charges itself the worst case up front: hourly rate times the hard TTL the instance will be killed at. That liability is written to the database before anything else happens, so a process that dies mid-lease doesn't take the bound with it. Elapsed time never enters the accounting.
On top of that sits a flat spend window: $10 per rolling day, counted from the first spend rather than from midnight. A calendar day can be gamed (the whole allowance at 23:59, the whole allowance again at 00:01), so the window starts when the money does.
Two smaller habits came with it. The cap is checked twice, before the search for an offer and again on the path that would commit the rent, because a query that can only end in a refusal is neither useful nor free. And every refusal writes one greppable line, GPU_RENT_REFUSED reason=…, so "why didn't it scale up?" is answered by the log rather than inferred from an absence. I say every. The next day I found one refusal path that still said nothing and fixed that too.
Then I deleted it
Five days after moving the fleet code in, I removed it: 70 files, 6,801 lines, 2,385 of them tests that were passing. Nothing had broken and nothing had been rented. The reason was plainer than a bug: inference had moved to the gateway, so nothing needed a GPU rented for it any more, and the loop's own README listed what it could do wrong (two autoscalers on one account destroy each other's instances; an instance whose label prefix changes bills forever). Infrastructure whose failure mode is permanently billing hardware should not sit one kubectl apply away from running.
It wasn't an rm -rf. Three things leaned on the fleet and each had to be answered.
- The proxy asked for a current lease and fell back to the CPU model without one. Only the autoscaler ever wrote a lease, so the lookup could now only return nothing. Unreachable by construction; it went.
- The concurrency provider summed capacity across the fleet. With no fleet, that's a constant, so the interface became one setting with the same default. No deployed value moved.
- Every replica published its load to a table every ten seconds, and the autoscaler was the only reader. Load still gates admission in-process; it just isn't written down for nobody any more.
What I kept, on purpose: the lease tables. No longer created or mapped, not dropped, because those rows are the only record of what the fleet ever spent. Deleting a spend record is a separate decision from deleting the code, and I left it separate.
The sweep afterwards found one more. The deployment manifest still passed --mode=gateway. The mode path was gone, so the argument bound to a key nothing reads. Inert, and it reads as though the service still had a mode, which is exactly the wrong thing for the next person to believe. Harmless to the machine, a lie to the reader.
If you build anything that pays per attempt, read the units on every limit you have. A timeout is seconds. A retry cap is attempts. A price cap is dollars per hour. If nothing on the list is plain dollars, you don't have a budget; you have a speed limit. Charge yourself the worst case when you create the lease, persist it before anything runs, and make every refusal say so out loud.