The Philosophical Developer — Chapter 54: Gold Without a Teacher
2026-09-19 · 5 min read

Two posts ago I shipped Decision-Maker and admitted the honest gap in it: the harness was real, the data was a bag-of-tokens toy, because that was the only way I could measure calibration. Last week I learned that flat answers were half a bug nobody had trained long enough. I keep circling the same line — a calibrated engine is only as honest as the gold you trained it on, and real gold is expensive. This is the post about where real gold comes from. Short version: we compile it from prose, and we never let one model judge another to do it.
The missing half
Decision-Maker scores candidates and returns probabilities. Training it needs a target: for every state, the probability each answer is right. That is the gold. The toy seed set had gold by construction — a known rule computed it. Point it at a skill, at a policy, at a runbook, and the rules are buried in paragraphs. Someone has to dig them out, and the tempting tool to do it is an LLM. Read the prose, emit a rule table. Easy.
I banned that. Not because it parses badly, but because of what it does to calibration.
Why a teacher model loses to a compiler
Ask a second model “is this rule faithful to the text?” and you get a judgment, not a proof. Judgments are biases wearing a distribution. They get distilled into the head and they quietly destroy the exact property I built the engine for. The literature calls this the teacher-student collapse; the practical version is a confident model that learned your rubric’s blind spot. Once a judge is in the loop, I can no longer claim the probabilities mean anything.
So the rule was hard: no model scores correctness. Gold has to be exact by construction.
That is why the compiler is structured the way it is. Everything a model can get wrong is checked by something that cannot be wrong:
- Extract reads prose and emits candidate rules. Recall-first, never trusted.
- Verify is the wall. Two checks, both constructive. First, every rule’s field has to resolve in the input schema and every outcome has to be a real member of the question’s criteria — a rule pointing at a field that does not exist is caught, not hidden. Second, a rule interpreter runs each rule over hand-marked states and asserts the branch it takes matches what a human anchored it to. No model judges; a deterministic interpreter and exact span matching do. The only human in the loop is marking a few trigger states per rule, and that is the point — that is a cheap, auditable judgment, not an opaque one.
- Synthesize enumerates every state the schema allows, renders each to prose, runs the same interpreter, and computes the soft gold. Two runs over the same input are byte-identical. I can diff the output bytes between runs and prove nothing moved.
Then the probe referee runs the skill’s own commands read-only and compares the measured world to the rules’ prediction. A mismatch flags a rule; it never silently reconciles.
The honest bit
It is deterministic, which is the whole point, but deterministic does not mean complete. A rule no anchor exercises stays incomplete for a human; nothing gets stamped verified that I cannot show. Extraction is model-dependent under the hood — temperature zero and a seed shrink the variance, they do not kill it; the verifier is what normalizes that, and it always has the last word.
The refusal is by design: prose with no pinnable answer — “make it sound natural” — gets exit code two and a no_gold tag, not invented confidence. If it cannot be calibrated, it does not get trained on.
Why Hermes is the first real win
Hermes runs on skills. Every skill is a SKILL.md: paragraphs of exactly the decision-dense prose this compiler was built to eat. “If the backend is down and the logs look clean, investigate before you blame the port.” That is a rule, sitting in prose, waiting to be compiled into calibrated gold.
The meaning of that is concrete. Hermes already has to make the same decisions over and over — which tool, retry or escalate, probe before fix. Today those are habits and heuristics with no measured confidence. Handed the compiler, they become a training set that knows what it does not know. An agent that returns “I don’t know, 0.62” and one that returns “I don’t know, 0.99” should act differently, and with compiled gold they can learn that margin from real prose instead of from a judge.
It generalizes past skills. A decision manual compiles to a router with p-gates. A config generator’s tier-picking compiles to exact gold, because the outcome is a discrete branch. Runtime ops gates — retry, rollback, escalate — are the best case, because they are observable, so the probe referee is at full strength there.
Concrete and checkable, all in the github.com/dark5un/decisionmaker repo: scripts/gold_compiler.py is the tool, research/compiler/ holds the corpus tables and the verifier’s anchors, docs/GOLD_COMPILER.md is the honest write-up, and scripts/history_to_gold.py is the plan-06 loader that sources gold from a historical log instead of compiled rules. make gold-compiler runs the whole CI gate — every corpus table verified, and a double-run that must come out byte-identical.
Let the limits be the wall, not the flattery. It is schema-bound, not skill-bound: a head stays calibrated only inside the schema it trained on. It makes a head more informed, not more general. And the business-history loader shows the source of gold is pluggable — when outcomes are measurable, history beats compiled rules, and you swap the loader instead of the harness.
The point is the same one from chapter 53, aimed now at the input side: do not make a model confident when the problem is that nothing checked the gold. Verify by construction, compute by construction, and leave the few real judgments to a person.