The Philosophical Developer — Chapter 53: When a Decision Engine Says "I Don't Know"
2026-09-19 · 4 min read

A week after I shipped Decision-Maker, I sat down to actually use it and got a wall of coin flips. Ask it to choose between rebasing and merging and it returned 49.9% / 50.1%. Ask it to rate the severity of a payouts outage and every bucket came back at a quarter. The engine was fast, healthy, and telling me nothing.
That flatness looked like a model that had not learned. It turned out to be two different problems wearing the same jacket — and the fix for one of them was not what I expected.
The head was never loaded
The first problem was embarrassing and simple. My serve entrypoint built a brand-new decision head and served that. Random weights, std 0.02, near-zero logits. Softmax of near-zero logits is a uniform distribution. The trained head — the whole point of the thing — was sitting in a checkpoint that nothing ever read.
A random head on top of a frozen backbone is the classic silent failure. The service is green, /health returns 200, requests flow, and every answer is a shrug. You only notice when you read the probabilities. The fix was one load_state_dict and a loud warning in the log when the checkpoint is missing.
That explained 199 of the 200 symptoms. Retrain proved the head loaded, and Q1 moved from a coin flip to a real decision with a real margin.
And then it was still flat
The second problem was subtler, and it is the one I want to remember.
This engine is trained by scoring every candidate and softmaxing — but the training target is the gold probability distribution, not a hard “this is the answer” label. That design gives you honest probabilities by construction. It also, quietly, gives you a head that is well-calibrated and flat. The literature calls the effect label smoothing, and it is not a bug; it is the price of measured confidence.
I measured it. The gold data for a four-way choice is itself soft — the best option only wins about 45% of the time on average. So even a perfect model caps around 0.28 mean confidence on choice. Flatness, part of the time, is the correct answer.
The field’s standard cure is post-hoc temperature scaling: divide the logits by a constant to sharpen the probabilities. I fitted one. It improved the log-likelihood against the soft target, and it blew up my boolean calibration from ECE 0.008 to 0.163. Sharpening made the model less honest, not more. That is the exact trade the calibration papers warn about, and it does not show up if you only look at one metric.
The real lever was neither temperature nor a fancier head. It was that I had trained for thirty-six steps. Three epochs on a two-hundred-row seed set is barely a warm-up. I retrained the same soft objective for ten times as long — three hundred and sixty steps — and the margin came back on its own. Choice peak went 0.27 → 0.35, score 0.54 → 0.71, all while ECE stayed around 0.06 on test and 0.045 on the held-out OOD split.
Hard labels, by the way, did not help. A hard-target head stayed flat on choice and nearly failed the calibration gate. The soft objective was right; I just had not trained it enough.
The whole thing, including the runs and the checkpoint, is at github.com/dark5un/decisionmaker.
The honest bit
The improvement is real and now it is serving. But the ceiling is the data, and the data is synthetic — word bags with known category priors, which is the only way I could measure calibration honestly at all. Point it at a real question — “should I rebase or merge?” — and low confidence is the correct answer, because that domain is not in the training set. More epochs sharpened a head that knew nothing about git, which is the difference between sharper and more informed.
Temperature sharpening is not the fix I hoped for this problem, and I will not ship it. A calibrated flat answer that tells you it does not know is worth more than a confident one that is wrong.
The lesson, stated plainly: verify that training output actually reaches inference, then check your step count before you blame the architecture, and do not make a model more confident when the problem is that it has not been trained long enough to have anything to be confident about.