The Philosophical Developer — Chapter 58: The Memory Budget Behind a Usable Local 125B
2026-10-08 · 11 min read

The question stopped being whether this model could start on my machine. It was which version I should run on each card, how much context I could ask it to keep, and whether a second model could stay useful alongside it.
The answer is yes: local inference on this 32 GB RTX 5090, 12 GB RTX 4070 Ti and 125 GB of system RAM is feasible, and, for my work, usable. It is not a free lunch. The quant and the context window move the costs around, and the one-million-token setting has a very different price depending on which GPU is reading it.
The model does not live in VRAM alone
One naming wrinkle first. Strata describes Qwen3.8-Flash-Next as a 125-billion-parameter model. The quantization repository lists 177 billion parameters because its count includes a 51.2-billion-parameter per-layer n-gram lookup table. Those are different accounting boundaries, not two competing model builds.[1][3]
This is a sparse mixture of experts: 512 routed experts per layer, with ten active for a token.[1] The model is much larger than either card’s VRAM, but inference does not require copying every byte into a GPU. Strata uses system RAM for much of the expert working set, VRAM for the GPU-resident weights and hot cache, and keeps the large lookup-table shard on the SSD.[2] In my configuration, KV streaming keeps 32K tokens resident on the GPU and holds the rest of the cache in host memory.

Conceptual layout, not a trace of every runtime memory transfer. The two GPU lanes can belong to separate services; the main model can also be split across both cards.
That architecture is the first choice we made: use the whole PC as the inference system, rather than treating the GPU’s VRAM as the only usable memory. It makes the model fit. It also means that a context window can consume system RAM and take a long time to prefill even when generation is still interactive.
The four ways to use the cards
There are four configurations that matter, and they answer different questions. “Both” means one Strata model split across the 5090 and 4070 Ti. The Coder is a separate, code-oriented model instance on the 4070 Ti, and can run alongside a 5090 main model. Strata’s Coder keeps 256 of the 512 experts, selected for code; it is a different capability trade, not just another quant of the general model.[2]

Top left: 5090 alone. Top right: 4070 Ti alone. Bottom left: one model split over both GPUs. Bottom right: two independent services, with the general model on the 5090 and Coder on the 4070 Ti.
The choice per scenario ended up like this:
| Scenario | Quant and context choices | Why |
|---|---|---|
| General model on the 5090 | IQ3_S at 256K, 524K or 1M | Keep the quant we already trusted for quality; make larger windows explicit opt-ins. |
| General model on the 4070 Ti | IQ3_XXS at 256K, 524K or 1M | A substantial speed improvement over IQ3_S, while staying close to its published task score. |
| General model split across both cards | IQ3_XXS at 256K | Better than IQ3_S on this two-card test, but still slower than the 5090 alone; no larger-context ladder was worth adding. |
| Coding beside a 5090 main model | Coder IQ1_M on the 4070 Ti at 256K, 524K or 1M | A separate specialist service, so coding work need not evict the main model from the 5090. |
The main Strata variants remain mutually exclusive: only one main-model arrangement runs at a time. Coder is the exception; it can coexist with a 5090 variant. The widget and arbiter make those choices explicit rather than silently moving a model or changing its context.
Why IQ3_S stayed on the 5090
The first benchmark question was whether a different quant could make the 4070 Ti more useful without giving up too much quality. I measured the candidates rather than assuming the name of a quant predicts performance on this model and this engine.
These are local Strata timings from one run per cell: 1,500 generated tokens and a 22,068-token prefill, with MTP drafting enabled, a warm page cache and the same engine configuration. They are measurements from this machine, not upstream leaderboard scores.
| Arrangement | Model / quant | Decode tok/s | 22K prefill tok/s |
|---|---|---|---|
| 5090 | General IQ3_S | 195.0 | 6,238 |
| 4070 Ti | General IQ3_S | 44.6 | 884 |
| 4070 Ti | General IQ3_XXS | 56.9 | 1,026 |
| 4070 Ti | General IQ2_XS | 70.3 | 1,228 |
| 4070 Ti | General Q2_0 | 67.2 | 1,297 |
| 4070 Ti | General UD-IQ4_XS | 29.5 | 578 |
| 4070 Ti | General UD-Q4_K_XL | 23.3 | 452 |
| 4070 Ti | Coder IQ1_M | 54.4 | 1,639 |
| Both cards | General IQ3_S | 157.5 | 1,972 |
| Both cards | General IQ3_XXS | 167.5 | 3,167 |
| Both cards | General IQ2_XS | 183.1 | 3,747 |
| Both cards | General Q2_0 | 179.0 | 4,055 |
| Both cards | General UD-IQ4_XS | 141.1 | 1,704 |
| Both cards | General UD-Q4_K_XL | 107.0 | 959 |
| Both cards | Coder IQ1_M, tested split | 155.3 | 4,510 |
Speed is only half the quant decision. On the publisher’s task-average test, IQ3_S scored 93.26 against 93.12 for BF16; IQ3_XXS scored 92.57, IQ2_XS 89.16 and Q2_0 89.07. These are ISTA-DASLab’s reported xhigh-reasoning results, not measurements I ran locally; the 0.14-point IQ3_S/BF16 difference is not evidence that the quant is better than its source model.[1][4]
That made the decisions fairly clear. IQ3_S stays on the 5090: it is the highest-scoring published option we tested and already the known-good choice there. IQ3_XXS goes on the 4070 Ti and in the two-card main-model variant: on the 4070 Ti it moves decode from 44.6 to 56.9 tok/s, a 28% gain, for a 0.69-point lower task average than IQ3_S. IQ2_XS and Q2_0 are faster in some parts of the 4070 Ti test, but they give up about four task-average points versus IQ3_S. They are reasonable speed-first alternatives, not the balance we chose.
The Unsloth UD-IQ4_XS and UD-Q4_K_XL files fit in the tested arrangements, but they were much slower here and we did not have a directly comparable published quality score for them. More bits in the filename did not make them the right operating point on this machine. Coder IQ1_M is a separate code-tuned, expert-pruned model; its scores should not be read as a general-purpose quality comparison.[2]
After applying the selected quant, separate one-off checks measured 52.6/1,023 tok/s on the 4070 Ti and 177/3,186 tok/s on both cards. These are later validation runs, not repetitions to average into the single-run comparison table.
Before the quant sweep, I upgraded Strata from 0.1.39 to 0.1.40.4 and measured the same model before and after. The 5090 moved from 185.1 to 187.7 decode tok/s, the 4070 Ti from 48.5 to 48.2, and the two-card setup from 166.0 to 178.3. That baseline helped isolate the later quant comparison: the IQ3_S-versus-IQ3_XXS rows above use the same engine build, rather than mixing an engine change into the quant decision.
The result for “both cards” was also a useful correction to intuition. Layer-splitting across two GPUs improved the IQ3_XXS figures over IQ3_S in that arrangement, but the 5090 alone on IQ3_S still decoded faster and prefills far more quickly. I kept the two-card option at 256K for cases where its combined memory footprint is useful; I did not add a larger-context ladder just because it was possible.
Context is a separate budget
The model’s trained context is 262,144 tokens. The larger windows use YaRN scaling: x2 for 524,288 and x4 for 1,048,576. That scaling is fixed when the engine starts, so each size needs its own Strata configuration and service variant. These extensions are experimental; passing a retrieval probe does not establish that every kind of reasoning remains reliable at one million tokens.

The rightmost pair represents the 5090 1M + Coder 1M test. The diagram is qualitative; the table has the measurements.
First, the 5090 ladder with IQ3_S. The needle test places two planted facts at different depths inside randomized filler, and asks for both after the full prompt has been read. It tests retrieval at depth, not long-document reasoning.
| Context | YaRN | RAM after full window | Decode tok/s | 22K prefill tok/s | Prompt read and retrieval |
|---|---|---|---|---|---|
| 256K | none | 64 GB | 194 | 6,160 | 252,914 tokens in 43 s; 2/2 |
| 524K | x2 | 67 GB | 189 | 6,210 | 516,890 tokens in 100 s; 2/2 |
| 1M | x4 | 74–75 GB | 176 | 6,022 | 1,023,221 tokens in 248 s; 2/2 |
The larger windows cost RAM and some decode speed, while the 5090 still read the million-token test prompt in 248 seconds, about four minutes. That is the reason 1M is an available option, not the default: it is useful when the task needs it, and wasteful for an ordinary chat.
The 4070 Ti and Coder variants went through their own near-full-window checks. “Short prefill” below is the same 22,068-token benchmark; the long-prompt timing is from the separate retrieval request.
| Instance | Context | Decode / short prefill tok/s | Long prompt | Read time / measured prefill | Retrieval |
|---|---|---|---|---|---|
| 4070 Ti IQ3_XXS | 524K | 49.7 / 781.2 | 516,890 tokens | 613.5 s / 842.6 tok/s | both checks passed |
| 4070 Ti IQ3_XXS | 1M | 50.9 / 537.3 | 1,023,221 tokens | 2,132.3 s / 479.9 tok/s | both checks passed |
| Coder IQ1_M | 524K | 46.6 / 1,281.8 | 516,890 tokens | 408.3 s / 1,265.9 tok/s | both checks passed |
| Coder IQ1_M | 1M | 49.2 / 770.5 | 1,023,221 tokens | 1,735.1 s / 589.7 tok/s | both checks passed |
The table is why I separate “can hold the context” from “pleasant to fill it.” The 4070 Ti and Coder did retrieve both facts at about one million tokens, but reading the prompt took 35.5 minutes and 28.9 minutes respectively. For ordinary conversations, the 256K option is the sensible starting point; 524K and 1M are deliberate long-document modes.
Two full windows, one tight memory margin
The last test paired the 5090 1M main model with Coder 1M on the 4070 Ti. The Coder window was filled first; then Strata read a 1,023,221-token prompt on the 5090, again retrieving both planted facts. With both windows full, the host reported 120,472 MiB used and 7,592 MiB available. The process completed without an out-of-memory failure, but that is not a lot of spare room for another large workload.
After both windows were full, the Coder measured 40.7 decode tok/s and 769.7 tok/s on the short prefill. Its solo 1M run measured 49.2 decode tok/s. The two services were shut down after the test, and host memory returned to its idle range. The result is not “run every large service together”; it is that this particular two-model arrangement fits, provided the rest of the machine stays quiet.

Conceptual view of the final test’s RAM pressure; the measurement was 120,472 MiB used and 7,592 MiB available.
That is the uncomfortable edge of the experiment: I was showing that the duo fits, not recommending it as a permanent near-the-limit operating point. In normal use I would leave one or both contexts shorter unless the task actually needs the full million-token windows.
The honest bit
These were engineering measurements, not a controlled paper. The quant sweep used one run per cell. The context probes used synthetic filler and two known retrieval targets; they do not test million-token reasoning, consistency across arbitrary long documents, or quality at every prompt depth. The long-prompt speeds are specific to this machine, this Strata build and these configurations.
There is also a human observation I want to keep separate from every table. To me, this model differs noticeably from its siblings and from other models below 125B: I need to interfere less in its decisions and actions. That is my experience, not a benchmark result.
What the measurements do establish is narrower and still useful: with IQ3_S on the 5090, IQ3_XXS where speed matters more, separate long-context presets, and a carefully pinned Coder instance, local inference is not just possible on this machine. It is something I can use, with known costs before I start it.
The series: The Philosophical Developer — real experiments, real code, real outcomes. Previous: Chapter 57: A Panel That Shows You the Plan Before You Click. Earlier in this arc: Chapter 56: The Lab Gets a Control Plane and Chapter 55: The Model That Outruns My Reading Speed.
Repo: ai-lab-quadlets — the systemd quadlets, arbiter, benchmark scripts and widget behind the lab.
Sources
[1] https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF — ISTA-DASLab Qwen3.8 Flash Next GSQ-RCO model card [2] https://github.com/Niko1221/Strata/blob/main/docs/MODELS.md — Strata model sizes, versions and memory guide [3] https://github.com/Niko1221/Strata/blob/main/README.md — Strata project README [4] https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/commit/c67535ccaa71f61547bb323a2828d5298221d83e — ISTA-DASLab IQ3_S release and benchmark results