The 388-Million-Parameter Laboratory
Why build what you can download?
The last year of language-model architecture has been a public conversation between labs. Qwen shipped the Gated DeltaNet — a recurrent layer that writes into an associative memory with a delta rule instead of attending to the past. DeepSeek made attention cheaper by compressing it into a latent space. Kimi answered with the Delta Attention: the same delta rule, but with a decay gate that operates channel by channel rather than per head. Each of these papers is a claim about how information should be forgotten. None of them comes with the thing that would let you check the claim yourself: a small, clean, hackable implementation.
So I built one. 388 million parameters, 20 layers, a 3:1 ratio of delta-net blocks to attention, one 24-gigabyte GPU. From v2 on there is also a 17M-parameter multi-token-prediction head; from v3 on, a 134M-parameter n-gram table. The headline number is the main path — the quantity the sweeps hold iso-parameter. It is not a toy — every block was cross-checked layout by layout against Qwen3.5's published weights and its reference implementation. It exists for one reason: at this scale, an architecture hypothesis costs 0.5 billion tokens and about 24 hours to test. The unit of work here is not the model. It is the claim.
Of Qwen3.5-0.8B's roughly 800M parameters, about 250M are the embedding matrix for a 248,320-token vocabulary. The first scaling decision was therefore not architectural at all: a 65,536-token in-house tokenizer cuts the parameter bill by ~185M before a single layer is designed. Shrinking the model is easy. Shrinking the hypothesis — keeping every mechanism the papers describe, at a size where a sweep arm fits in a day — is the actual craft.
Four versions, one change each
The laboratory only functions because each version changes exactly one thing, guarded by a test that would catch a second. The lineage, so far:
The guarantee that makes the lineage scientific lives in a single test. v4's gate is a low-rank projection producing 1,536 decay values per layer where v2 produced 12 — a structurally different object. But if you copy the old per-head weights into the low-rank factors and repeat them across channels, the new model must reproduce the old one exactly. It does:
Architecture comparisons die from confounds: a new gate and a wider FFN, a new kernel and lower precision, a memory module and a bigger table than either arm can fairly afford. v4's parameter bill was balanced to 387,539,328 against v2's 387,677,928 — a −0.036% difference, under the 0.5% tolerance the test suite enforces. And because the head-wise reconstruction matches to 2.4e-07, any difference the sweep measures can only come from the gate. If the result is real, it is attributable; if it is noise, that is measurable too.
Launch day: two lies and a bit-exact fix
v4 began training on the evening of September 13. The launch log, kept verbatim, is worth showing because it contains two mistakes and one fix — in that order, as they happened:
My "13.2 GB peak" claim was wrong — I sampled at 8-second intervals at one moment; the true peak of the micro-run could be higher. A number taken at the wrong moment is not a measurement.
The whole episode cost about two hours. The cheaper lesson: a micro-benchmark is only evidence about the model you actually loaded, and a memory figure is only evidence at the moment the allocator peaked. Both mistakes were caught by checking, not by luck.
What the system does
Every block was validated against published weights and the reference implementation before any experiment touched it. Innovation on an unverified base is indistinguishable from debugging.
Each version ships exactly one architectural delta, guarded by an identity test and iso-parameter accounting, evaluated against the previous version on paired held-out data. The current tally: v3's n-gram memory bought −0.0375 nats per token on average across the 10 paired checkpoints of steps 1000–1999 (final held-out CE 1.6248 vs 1.6609; minimum gap 0.018 against a ±0.01 noise band) and +7.5/+8.5 points of exact-match on needle retrieval — with token-copy unchanged, proving the gain is specific, not general progress.
The reading rules for the running sweep were fixed before launch: the noise band is ±0.01 nat, two seeds minimum per arm, and — the important one — a null result at 2,000 steps will not be read as contradicting Kimi, because the channel-wise gate's demonstrated advantage is long-context memory control, and the sweep's context is 1,024 tokens. The long-context suite already shows the ceiling: beyond the training length, retrieval holds at depth 0.9 and collapses at depth ≤0.5.
Δ ≈ 0 at 2000 steps / seq 1024 does NOT contradict Kimi — the benefit of the channel-wise gate is long-context. The annealed run at 4096 is the real judge. — HANDOFF-v4.md, reading criteria, fixed September 13
Timeline
Gated-DeltaNet reimplemented from scratch, verified against Qwen3.5 GGUF layouts and the transformers fallback kernels. CE at initialization lands on ln(vocab) to two decimals — the symptom that had summarized every earlier init bug.
25% less associative state pays for a 33% wider FFN at iso-parameters, plus causal convolutions on attention q/k/v, residual-variance-scaled init, and an MTP head. The first controlled sweep compares v2 against v1 on equal footing.
Engram-style PLE: 134M-parameter hash table, gated injection before block 2. Sweep result: −0.0375 nats/token on average over the paired checkpoints of steps 1000–1999, 10/10 in v3's favor (final CE 1.6248 vs 1.6609); long-context eval shows +7.5/+8.5 exact-match on needle/multi-key at 0.52B tokens — reaching v1-at-1.96B performance on multi-key retrieval at a quarter of the training compute.
Channel-wise decay gate, low-rank projection, kernel rewritten (chunk 32, einsum decay ratios), gradient-checkpointed bit-exactly after the OOM. Two new arms — channel-wise gate with and without the n-gram memory — two seeds each, four runs in total: roughly four days at ~6,000 tok/s per run, with a model checkpoint every 1,000 steps and automated status checks every six hours.
On this machine, that is what reproducibility means.
The next article in this series will be the v4 verdict — channel-wise against scalar decay, and the interaction table against the n-gram memory — followed, for whichever arm wins, by annealed training at 4,096 tokens and the long-context suite out to 4,096. That last number is the one the current architecture cannot reach: retrieval holds at the end of context and collapses in the middle. The channel-wise gate is, among other things, a hypothesis about exactly that failure.
This article was written while the v4 sweep was running on the author's machine (a single RTX 3090). All v1-v3 figures are final, from the sweep logs of September 3; all v4 figures quoted are infrastructure measurements (memory, throughput), not results. Architecture and test design were developed in collaboration with Kimi (Moonshot AI); training, sweeps and failures are real, on real hardware, with timestamps. Version provenance, in short: v1 and the v3 n-gram memory were dax's ideas — implemented respectively by a local 27B running on the training machine, and by Kimi K3; v2 was designed and implemented by Fable 5; v4 was implemented by Kimi K2.8 preview from the Kimi Linear architecture. The sweep protocol, the reading criteria and the final judgment are the author's alone. The reading criteria quoted above were fixed in writing before the v4 launch and will not be revised after the results arrive.
REFERENCES
- Gated Delta Networks: Improving Mamba2 with Delta Rule — the recurrent operator the whole lineage is built on.
- Kimi Linear: An Expressive, Efficient Attention Architecture — the channel-wise decay gate v4 is testing at 388M.
- Engram — the sparse n-gram memory behind v3's PLE.
- Qwen3.5 model weights (GGUF) and the transformers reference implementation — the verification target for v1.
Next: Behind the Curtain — how a language model is built from scratch · ← back to blog