The Philosophical Developer — Decision-Maker: Calibrated Probabilities, Not Parsed Text
2026-09-18 · 4 min read

Ask an LLM a yes/no question and the honest answer is a wall of text you then have to parse. “Which team should handle this?” comes back as a paragraph, and somewhere in it there is an answer you trust differently than the sentence before it. That never sat right with me. So I built the opposite: a local service that takes a state and a set of typed questions and returns a real probability distribution — calibrated, not hopeful.
This is Decision-Maker. It lives mostly in my office, speaking HTTP on port 8090.
The idea
The core move is to stop treating a question as text generation and start treating it as scoring. Every candidate answer becomes a leaf — a small prompt ending in “Decision:” — and all of them are packed into a single forward pass through a frozen backbone. A small trained head then reads the hidden state at each leaf’s last token and emits one scalar per candidate. Softmax over a question’s candidates is the answer.
There is no token decoding anywhere. The output_tokens count is always zero. The model never says a word; it just scores. That single fact is responsible for most of the speed and all of the reproducibility.
Three question types map to three answer shapes:
- boolean — a probability of “yes”, one float, no confidence field
- choice — one winner plus a probability per option that sums to one
- score — a position on an ordered scale, fractional, with per-level probabilities
The calibrated part
Scores are not opinions; they are a distribution, and a distribution can be measured. The head is fine-tuned on a proper-scoring objective — cross-entropy as the primary arm, Brier as the strongest one on the seed run — plus an RLCD-style paired proper-reward arm that is the experiment, the one I honestly expected might not win. On the seed set it did not beat plain CE on test. That was a research outcome, not a bug; I reported it as such.
The gate is boolean ECE over ten fixed bins, on test and on a deliberately held-out OOD split:
| loss | test ECE | OOD ECE | gate |
|---|---|---|---|
| ce | 0.0261 | 0.0424 | PASS |
| brier | 0.0002 | 0.0147 | PASS |
| paired | 0.0436 | 0.0247 | PASS |
confidence on the wire is not “how likely this is right.” It is how peaked the returned distribution is on its winner — a gate for act / review / escalate. Boolean carries no confidence at all, because a confident “no” and a confident “yes” are equally confident.
The contract
The API is my own, and deliberately so. A request is {state, questions} where each question carries a type and criteria. Answers come back under the same qids, and the qid is never part of inference — rename a question and the answer distribution does not move. Errors are typed, never a bare 500: validation, internal, not_ready. There is a pinned model and a strict revision, so the answer for a given input is stable unless I change it on purpose.
I wrote thin typed clients in Go and in Rust, and then spent a session proving they agree with each other and with a raw request, to nine decimal places. That is the kind of check that sounds like overkill until it catches a real disagreement.
The container bit
It runs under rootless Podman as a quadlet, pinned to a GPU by UUID — never by device index, because those minors flip across reboots and two tools can number the same two cards in opposite orders. I have burned myself on that exact trap more than once. The engine refuses to claim GPU success without a real device present. Load once, warm, serve.
The code is at github.com/dark5un/decisionmaker, if you want to poke at it.
The honest bit
The data here is synthetic — a bag-of-tokens task with known category priors, which is how I could measure calibration honestly at all. The moat is meant to be real domain data in the same schema, with a distribution as the target rather than just a label. That part is not done; the harness is, the data is not.
The ECE numbers understate themselves: predictions concentrate near 0.4, so most fixed bins sit empty and the weighted ECE adds zero for them. Use the arms to rank, not the raw figures. And one draw per arm means a fresh seed would sharpen the ordering.
None of this is magic. It is a frozen backbone, a head with a LayerNorm and a linear layer, and an honest measurement of whether the probabilities match the outcomes. The point is that it is calibrated by construction, repeatable by contract, and small enough to run in my own office.