The Philosophical Developer — Chapter 60: The Harness Tab Strata Never Had
2026-10-09 · 6 min read

Strata runs a 125B model on one consumer GPU and ships a web UI I genuinely like: three tabs — Chat, Monitor, About — dark, token-driven, no build step, no framework. DeepSeek’s harness (dsh) is the other thing on this machine: a full agent loop with sessions, approvals, compaction, subagents. The obvious move was to run dsh next to Strata and point it at the same model. I did that first, and it stayed a second app with a second session store, guessing at what the inference server was doing.
The less obvious move, and this chapter’s work: give Strata a fourth tab. Same design system, same server process, same monitor numbers — and an agent loop built natively for Strata’s architecture. Nine milestones, one fork, and a rule I keep proving: the guardrails are the feature.
What “native” actually bought
The harness is a Python package inside Strata’s serve layer — serve/harness/ — and the tab is a thin client: an SSE event stream plus POST commands, exactly how the Chat tab already streams. That placement is the whole design:
- The agent loop talks to the engine through the same in-process Service the web UI uses. No second HTTP hop, no second model slot, no guessing at context pressure — the loop reads the monitor’s own metrics and compacts when the window fills.
- Sessions are an append-only JSONL event log. The LLM’s message history is derived from a projection of that log; compaction shadows spans instead of rewriting them. Reload the browser mid-turn and the session resumes from the log; fork it and the child inherits the same durable history.
- The fork stays upstreamable by construction: 14 added lines in
server.py, zero deleted, two route branches and one wiring call. Everything else lives behind that boundary.
The honest bit: none of this is faster than running stock dsh. It is legible. Every model request is reconstructable from the log (the request/header event carries the full derived message set, sampling, and a hash of the tool schemas), every tool call has its approval row, every retry has an attempt event. When an agent session goes sideways, I read the log, not the UI.
Approvals are a UX problem, not a checkbox
The first live session on the 5090 delivered working code — 11 passing tests, a real deliverable — and the transcript made it look like garbage. Reasoning was captured from the engine and never painted, so tool calls appeared unmotivated. Approval previews truncated mid-command: the user saw ... && echo "=== LINT ===" && and nothing after. You cannot approve what you cannot read.
So the milestone loop kept pulling back to the rendered page: a standing rule that every UI-touching change ends with a screenshot at a viewport matrix, geometry measured with getBoundingClientRect(), and a vision pass on the shots. The rules that came out of it:
- A pending approval is a loud card with the FULL payload — the whole command, the path as the hero line for writes, previews cut at line boundaries.
- A decided approval folds to one quiet line. History must not compete with the live step.
- Content labels carry the payload kind and size: “Result · tabular output (325 chars, 12 lines)”. The scan layer answers “worth opening?”; the label answers “what am I looking at?”
Yolo mode got the same treatment: auto-allow is per-session, never survives a restart, shows a loud banner, and every unattended run still logs a decided row. Fail-closed everywhere: a missing answerer is “unavailable”, never an implicit allow.
The plugin question, deferred until it was safe
“Does the harness support plugins?” took three milestones to answer. First MCP: Strata already speaks MCP for the Chat tab, so the harness registry attaches the same hub’s tools as namespaced, always-gated entries — an out-of-process server can do anything the machine allows, so every call asks. Then the loop hook surface: pre-step, pre-tool, post-tool observer lists, so policies become plugins instead of loop branches.
Self-authored plugins stayed deferred the longest, because a self-written plugin is arbitrary Python that auto-loads on every future start — one approved write_file becomes permanent code execution with no approval card. When the decision finally went the other way, it shipped as two independent gates:
- Load review: a plugin file is imported only when its sha256 matches a reviewed manifest. New or changed files are quarantined — listed, never imported. The review UI shows the full source and a diff against the last reviewed snapshot. Approval is per exact bytes, so editing a plugin re-quarantines it.
- Write gate: writes into the plugin directory are never allowlistable. Even under yolo, that write forces the loud card, and the card says what it is: “This file becomes server code after you review it in the Plugins panel. It will NOT load automatically.”
Neither gate is bypassable by the other. That is the whole security model in two sentences.
The gate that caught a silent fork bug
Rebase discipline has a teeth version. After re-basing onto 126 commits of upstream drift, the harness gate stayed green while 22 serve-suite tests failed: one of my own commits had silently replaced server.py with a pre-rebase copy, dropping ~309 lines of upstream code — a watchdog, a body-limit fix, slot handling. The harness’s own tests could not see deleted upstream code because they never run it.
The rule since: after any rebase, run the full suite, not just the feature gate. 750 serve tests, 180 harness tests, ruff, pyright, eslint — every change, every slice tip. The fork ships as eight stacked PR-sized branches (pr/1-scaffold through pr/8-plugins), each gate-green at its own tip, so the series can be offered upstream slice by slice if the day comes.
Small footprint, real model
The last check before trusting it for daily work: does the fork run on the small card? Inside the same container image Strata ships, with the repo mounted and the entrypoint overridden to the fork’s server: IQ3_XXS, 64K context, int8 KV — 11.6 of 12 GB on a 4070 Ti. A real session over HTTP: reasoning painted, two shell approvals answered from the tab, tool results, a model-generated session title. The harness process itself is light; the model is the footprint.
The honest bit: with both big models loaded, host RAM got tight — this box runs one large model at a time, and the harness rides whichever one is up.
What it is now
A Strata with a Harness tab: multi-turn agent loop, sessions that survive reloads and forks, compaction driven by the monitor’s own numbers, subagents with budget caps, presets as YAML, MCP tools behind approval cards, plugins behind two gates. Same look and feel, same single process, same loopback default.
It is an experiment with an unknown ceiling, kept honest by logs, tests, and a viewport matrix. The tab is open at localhost, and for the first time the agent and the inference server are reading the same dashboard.