The Philosophical Developer — Chapter 55: The Model That Outruns My Reading Speed

2026-10-05 · 6 min read

The Model That Outruns My Reading Speed

I have not seen 160 tokens per second before. Not on a server rack, not on a cloud GPU I was renting by the minute, not in a benchmark table I was meant to believe. And now it is happening in my house, on a gaming card, feeding the agent harness I use to write these chapters — and the harness’s output arrives faster than I can read it. I have to scroll to keep up with my own toolchain.

This chapter is about Strata: an open-source engine that runs a 125-billion-parameter model on a normal PC, and what changed in my daily work the day it became the brain of the lab.


What Strata actually is

Strata is an engine by Niko1221 that runs Qwen3.8-Flash-Next — a 125B-parameter model that normally wants a server — on a gaming PC. Twelve gigabytes of VRAM or more, 32 GB of RAM, about 80 GB of disk, and the model lives entirely on your machine. Nothing leaves the box. The README’s own measurements put it at 53 tokens per second on an RTX 5070 for the IQ3_S quant, with an honest note that a 24 GB card “should write about 100-140 tokens per second.”

I read that line and filed it under we’ll see.

My card is an RTX 5090 with 32 GB of VRAM. The model I serve is qwen3.8-flash-next-iq3_s — roughly 84 GB of IQ3_S weights, which spills into system RAM by design; Strata’s whole trick is keeping the hot layers on the GPU and letting big RAM cover the rest. I run it with a 262,144-token context.

The quadlet, because of course it’s a quadlet

Chapter 45 packaged the lab; this is the lab growing a new organ. Strata joins the stack as one more Podman Quadlet — systemd-strata, pinned to the RTX 5090 by GPU UUID, built for CUDA architecture 120 from the upstream Dockerfile, published on 0.0.0.0:11434/v1 with an API key required on every request. The installer (scripts/install-strata.sh in ai-lab-quadlets) builds the image, generates a 64-character key, writes the unit, and then — deliberately — does not start it. The first start is the ~84 GB model download, and that decision belongs to me, not to an installer.

One port detail worth recording: Strata moved from 11437 to 11434, the Ollama-default slot, to sit with the other model APIs in the 1143x family. The registry in services.json is the only writer for ports now — more on that discipline in the next chapter.

The numbers I actually measured

Not the vendor’s table. Mine, today, against the live service:

$ curl -s http://127.0.0.1:11434/health
{"status":"ok","model":"qwen3.8-flash-next-iq3_s","max_context":262144,
 "api_key":true,"loaded":true,"service":"strata"}

A short reply came back at 135.8 tokens per second (7.36 ms/token), with speculative decoding doing real work — 34 draft tokens proposed, 25 accepted. A 547-token completion arrived in 3.5 seconds wall-clock, which is about 158 tokens per second end to end. That is the number my eyes keep failing to catch up with. The vendor’s projection for a 24 GB card was 100-140; the 32 GB card lands at the top of that band and above it.

The honest bit: prompt processing on tiny prompts measured only ~213 tokens/s, because short prompts are dominated by overhead — the big prefill numbers in the vendor tables (1,600+ tokens/s) need a real document in front of them. And 262K context is not free: the KV cache raises memory use, and I watched RAM and VRAM on the first long session like a hawk. Speculative decoding is doing a third of the work in those decode numbers; it is not magic, it is acceptance rates.

What it changed in the harness

Here is the meta bit: this chapter is being written through the harness that Strata powers. The agent that steers my sessions — the one reading repos, running tests, drafting these posts — now thinks at 135+ tokens per second on a local model with a quarter-million-token context.

Three things changed that I did not expect:

The loop stopped feeling like waiting. At 20-40 tokens/s you learn to fire a prompt and walk away. At 135+ you stay in the chair. The agent’s reasoning trace streams faster than I can skim it, which sounds like a problem and is actually a discipline: I read the decision, not the monologue, and I intervene on the plan instead of the prose.

Context stopped being a budget. 262K tokens means the harness carries the whole repo, the session history, and the reference docs without me rationing what stays in the window. The methodology chapters about traceability were built for a world where I had to prune; that world quietly ended.

The privacy line moved. The model that reads my company’s repos, my client-adjacent notes, and my invoices is a process on my own machine behind an API key. The only network traffic is the model download.

It is not the smartest model in the world. It is a 125B quant, and it loses to frontier APIs on the hardest reasoning. But it is mine, it answers in milliseconds, and at this speed the harness stops being a turn-based conversation and starts being a live pair programmer.

Try it

If you have 12+ GB of VRAM and 32+ GB of RAM, Strata’s README is a genuinely good install guide — it even has an AI_SETUP.md page designed to be pasted into your coding agent, which is its own quiet joke about where we are in 2026. If you want it the way I run it — rootless, quadlet-pinned, key-gated, boot-persistent — the ai-lab-quadlets installer builds and wires it in one command.

I kept saying I would believe 160 tokens per second when I saw it. I have not seen 160. I have seen 158, and my reading speed is now the bottleneck.


The series: The Philosophical Developer — real experiments, real code, real outcomes. Previous: Chapter 54: Gold Without a Teacher. Next: Chapter 56: The Lab Gets A Control Plane.

Repos: