DAXZEIT
SEPTEMBER 2026 · DAXZEIT · co-authored with Kimi K2.8 Preview

The 388-Million-Parameter Laboratory

Rebuilding Qwen's hybrid delta architecture — and now testing Kimi's — on a single 3090, one verified change at a time.

Why build what you can download?

The last year of language-model architecture has been a public conversation between labs. Qwen shipped the Gated DeltaNet — a recurrent layer that writes into an associative memory with a delta rule instead of attending to the past. DeepSeek made attention cheaper by compressing it into a latent space. Kimi answered with the Delta Attention: the same delta rule, but with a decay gate that operates channel by channel rather than per head. Each of these papers is a claim about how information should be forgotten. None of them comes with the thing that would let you check the claim yourself: a small, clean, hackable implementation.

So I built one. 388 million parameters, 20 layers, a 3:1 ratio of delta-net blocks to attention, one 24-gigabyte GPU. From v2 on there is also a 17M-parameter multi-token-prediction head; from v3 on, a 134M-parameter n-gram table. The headline number is the main path — the quantity the sweeps hold iso-parameter. It is not a toy — every block was cross-checked layout by layout against Qwen3.5's published weights and its reference implementation. It exists for one reason: at this scale, an architecture hypothesis costs 0.5 billion tokens and about 24 hours to test. The unit of work here is not the model. It is the claim.

model.py, config header — August 2026 """Scaling 300M vs Qwen3.5 0.8B : hidden=1024 (identique), FFN=2048 (réduit de 3584), depth=20 (réduit de 25), vocab=65536 (réduit de 248320 — le vocab de Qwen3.5 coûte ~250M params, c'est un tiers du modèle, pas de l'architecture)."""
The vocabulary is a third of the model

Of Qwen3.5-0.8B's roughly 800M parameters, about 250M are the embedding matrix for a 248,320-token vocabulary. The first scaling decision was therefore not architectural at all: a 65,536-token in-house tokenizer cuts the parameter bill by ~185M before a single layer is designed. Shrinking the model is easy. Shrinking the hypothesis — keeping every mechanism the papers describe, at a size where a sweep arm fits in a day — is the actual craft.

Four versions, one change each

The laboratory only functions because each version changes exactly one thing, guarded by a test that would catch a second. The lineage, so far:

v1 — Replicate. A from-scratch Gated-DeltaNet, verified block by block against the Qwen3.5 GGUF (shapes, dispositions, config) and the transformers reference. The slow, plain-PyTorch delta rule is the learning path; fast kernels are explicitly deferred.
PROVENANCE — idea: dax · code: dax's local 27B, running on this same 3090
v2 — Reallocate. Same operator, different budget: 25% less associative memory (12 instead of 16 state heads) pays for a 33% wider FFN, at strictly iso-parameters — plus a multi-token-prediction head, trained only.
PROVENANCE — idea & code: Fable 5 (Anthropic)
v3 — Remember n-grams. An Engram-style hash table addressed by local 2- and 3-grams, read once per token, gated by the hidden state, injected before block 2. The sparse memory from the Qwen3.8-Flash-Next recipe, at 35% of the main path's parameters.
PROVENANCE — idea: dax · code: Kimi K3 (Moonshot AI)
v4 — Forget precisely. The Kimi move: the decay gate goes from one scalar per head to one value per state dimension, projected low-rank from the hidden state. Currently training as this article is published.
PROVENANCE — code: Kimi K2.8 preview (Moonshot AI), from the Kimi Linear / Qwen3.5 architecture papers

The guarantee that makes the lineage scientific lives in a single test. v4's gate is a low-rank projection producing 1,536 decay values per layer where v2 produced 12 — a structurally different object. But if you copy the old per-head weights into the low-rank factors and repeat them across channels, the new model must reproduce the old one exactly. It does:

test_model_v4.py — identity check, September 13, 2026 test_v4_matches_v2_with_headwise_gate v4 α head-wise == v2 : écart max 2.38e-07
A test that catches a second variable

Architecture comparisons die from confounds: a new gate and a wider FFN, a new kernel and lower precision, a memory module and a bigger table than either arm can fairly afford. v4's parameter bill was balanced to 387,539,328 against v2's 387,677,928 — a −0.036% difference, under the 0.5% tolerance the test suite enforces. And because the head-wise reconstruction matches to 2.4e-07, any difference the sweep measures can only come from the gate. If the result is real, it is attributable; if it is noise, that is measurable too.

Launch day: two lies and a bit-exact fix

v4 began training on the evening of September 13. The launch log, kept verbatim, is worth showing because it contains two mistakes and one fix — in that order, as they happened:

The micro-run measured the wrong model. The first memory and throughput figures (20.9 GB, ~9,000 tok/s) looked fine — but they came from v1. The quick-launch script hardcodes v1; only the architecture-aware launcher points to v4. The benchmark had run the baseline, not the candidate.
The real model OOMs at batch 4. With the true v4, batch 4 peaked at 22.9 GB and batch 2 at 21.5 GB. The cause is structural: with a channel-wise gate, the decay-ratio tensor γ is (batch, heads, chunk, chunk, channels) in fp32, saved for the backward pass. It does not shrink proportionally with the batch size the way an attention matrix does.
The fix is verified bit-for-bit. Gradient-checkpoint the delta-rule kernel in training mode: recompute the γ chain during backward instead of saving it. Verified diff 0.0 against the un-checkpointed forward before being trusted — a bit-exact fix is the only kind you don't have to re-validate. Relaunched: 20.8 GB steady, 98% GPU utilization, ~6,000 tok/s, about 24 hours per arm.
sweep-v4-wrapper.log — first checkpoint after relaunch, September 13-14, 2026 pas 0 | loss 14.6565 | lr 1.50e-07 | gnorm 4.99 | 773 tok/s | 0.000 B | 0.09 h pas 10 | loss 14.6409 | lr 1.65e-06 | gnorm 4.93 | 3737 tok/s | 0.003 B | 0.21 h pas 20 | loss 14.5862 | lr 3.15e-06 | gnorm 6.66 | 4566 tok/s | 0.006 B | 0.33 h 20739 MiB, 98 %
My "13.2 GB peak" claim was wrong — I sampled at 8-second intervals at one moment; the true peak of the micro-run could be higher. A number taken at the wrong moment is not a measurement.

The whole episode cost about two hours. The cheaper lesson: a micro-benchmark is only evidence about the model you actually loaded, and a memory figure is only evidence at the moment the allocator peaked. Both mistakes were caught by checking, not by luck.

What the system does

It replicates before it innovates

Every block was validated against published weights and the reference implementation before any experiment touched it. Innovation on an unverified base is indistinguishable from debugging.

It changes one thing at a time

Each version ships exactly one architectural delta, guarded by an identity test and iso-parameter accounting, evaluated against the previous version on paired held-out data. The current tally: v3's n-gram memory bought −0.0375 nats per token on average across the 10 paired checkpoints of steps 1000–1999 (final held-out CE 1.6248 vs 1.6609; minimum gap 0.018 against a ±0.01 noise band) and +7.5/+8.5 points of exact-match on needle retrieval — with token-copy unchanged, proving the gain is specific, not general progress.

It writes its verdicts in advance

The reading rules for the running sweep were fixed before launch: the noise band is ±0.01 nat, two seeds minimum per arm, and — the important one — a null result at 2,000 steps will not be read as contradicting Kimi, because the channel-wise gate's demonstrated advantage is long-context memory control, and the sweep's context is 1,024 tokens. The long-context suite already shows the ceiling: beyond the training length, retrieval holds at depth 0.9 and collapses at depth ≤0.5.

Δ ≈ 0 at 2000 steps / seq 1024 does NOT contradict Kimi — the benefit of the channel-wise gate is long-context. The annealed run at 4096 is the real judge. — HANDOFF-v4.md, reading criteria, fixed September 13

Timeline

AUGUST 2026 — v1: the reference

Gated-DeltaNet reimplemented from scratch, verified against Qwen3.5 GGUF layouts and the transformers fallback kernels. CE at initialization lands on ln(vocab) to two decimals — the symptom that had summarized every earlier init bug.

AUGUST 2026 — v2: the reallocation bet

25% less associative state pays for a 33% wider FFN at iso-parameters, plus causal convolutions on attention q/k/v, residual-variance-scaled init, and an MTP head. The first controlled sweep compares v2 against v1 on equal footing.

SEPTEMBER 3, 2026 — v3: the n-gram memory

Engram-style PLE: 134M-parameter hash table, gated injection before block 2. Sweep result: −0.0375 nats/token on average over the paired checkpoints of steps 1000–1999, 10/10 in v3's favor (final CE 1.6248 vs 1.6609); long-context eval shows +7.5/+8.5 exact-match on needle/multi-key at 0.52B tokens — reaching v1-at-1.96B performance on multi-key retrieval at a quarter of the training compute.

SEPTEMBER 13, 2026 — v4: the Kimi experiment, launched

Channel-wise decay gate, low-rank projection, kernel rewritten (chunk 32, einsum decay ratios), gradient-checkpointed bit-exactly after the OOM. Two new arms — channel-wise gate with and without the n-gram memory — two seeds each, four runs in total: roughly four days at ~6,000 tok/s per run, with a model checkpoint every 1,000 steps and automated status checks every six hours.

The verdict for v4 is already written down, two days before the last token is seen.
On this machine, that is what reproducibility means.

The next article in this series will be the v4 verdict — channel-wise against scalar decay, and the interaction table against the n-gram memory — followed, for whichever arm wins, by annealed training at 4,096 tokens and the long-context suite out to 4,096. That last number is the one the current architecture cannot reach: retrieval holds at the end of context and collapses in the middle. The channel-wise gate is, among other things, a hypothesis about exactly that failure.

The specimens in this article are verbatim from model.py, test_model_v4.py, sweep-v4-wrapper.log and HANDOFF-v4.md, September 2026. The v4 sweep is running as of publication; no v4 result appears in this article, by design.
Methodology

This article was written while the v4 sweep was running on the author's machine (a single RTX 3090). All v1-v3 figures are final, from the sweep logs of September 3; all v4 figures quoted are infrastructure measurements (memory, throughput), not results. Architecture and test design were developed in collaboration with Kimi (Moonshot AI); training, sweeps and failures are real, on real hardware, with timestamps. Version provenance, in short: v1 and the v3 n-gram memory were dax's ideas — implemented respectively by a local 27B running on the training machine, and by Kimi K3; v2 was designed and implemented by Fable 5; v4 was implemented by Kimi K2.8 preview from the Kimi Linear architecture. The sweep protocol, the reading criteria and the final judgment are the author's alone. The reading criteria quoted above were fixed in writing before the v4 launch and will not be revised after the results arrive.

REFERENCES

This article opens a series on the 388M.

Next: Behind the Curtain — how a language model is built from scratch · ← back to blog