Behind the Curtain: What Happens When You Build a Language Model From Scratch?
You already know what it does. Do you know where it comes from?
You type a question, it answers. Somewhere in that reflex there is an unexamined assumption: that a language model is a thing — downloaded, installed, talked to. It is also a process, and the process is almost never shown. There are a thousand articles explaining how transformers work. This is not one of them. This is about how you make one: the five ingredients — data, tokenizer, architecture, a training loop, and judgment — followed through a real build, on a real machine, with the failures left in.
The machine is one RTX 3090, a consumer GPU with 24 gigabytes of memory. The model is a 388-million-parameter language model, built from nothing over four versions across six weeks. Nothing here requires a datacenter. That is the point.
It starts with text, not GPUs
The first ingredient is the least glamorous: a large pile of text, cut into tokens — the pieces a model actually reads. Not words. A word like “unbelievable” might be three tokens; a common word is one; a space can be part of one. The model will never see letters or words. It will see tokens, ten digits from now and forever.
The first corpus held 1.034 billion tokens — thirty programming languages, ten thousand documents each, plus some prose. Training began, and the curve did something annoying: it flattened. After a billion tokens, the model’s held-out score — its exam grade on text it had never seen — stalled at 1.6453 and crawled. Two explanations fit the curve perfectly, and the curve alone cannot tell them apart: either the model was full — out of capacity — or it was hungry, out of data.
The training pipeline reads the corpus by drawing random windows, with replacement. Dividing tokens consumed by corpus size gave the answer: 0.97 epochs. The model had read almost every document essentially once. It was not saturating its capacity. It was saturating its corpus — two conditions indistinguishable on the loss curve, separated only by counting how much data was left.
The fix was more world: a large set of educational prose, and three programming languages chosen deliberately — Python, shell, TypeScript — added and deduplicated (95,294 duplicate documents caught and dropped). The corpus grew ×3.49 to 3.611 billion tokens, and the mixture shifted from 82% code to 73%.
Then the measurement that makes the story worth telling. Training resumed from the same weights, against the new data — the cleanest experiment this project ever ran, because nothing changed except the world the model was reading. The improvement rate doubled, and stayed doubled:
One cautionary footnote lives inside that table. The very first eval after the swap showed an apparent leap — 1.6453 to 1.6141 in 85 steps — double the true rate. It was a transient, and the rule drawn from it now governs every conclusion this laboratory publishes: never conclude from a single point. Each evaluation carries 98,304 tokens of evidence; that is a sample, not a verdict.
The first real decision is the alphabet
Before a single layer is designed, one number quietly sets the budget: the size of the vocabulary. The reference model this project is scaled from — Qwen3.5-0.8B — spends about 250M of its ~800M parameters on the embedding matrix for a 248,320-token vocabulary. A third of the model is spelling.
So the first scaling decision was to cut the alphabet: a 65,536-token in-house tokenizer, trained on the corpus itself. That single choice frees roughly 185 million parameters before architecture begins — budget that can go to thinking instead of spelling.
After the corpus swap, compression improved from 3.18 to 3.43 bytes per token — with the tokenizer untouched. The old corpus spread thirty languages thin, so the tokenizer’s merges for any one of them were undertrained. Concentrating on three languages finally put the right merges to work. A model and its alphabet shape each other in both directions.
Then: how should it think?
With the alphabet chosen comes the question the architecture papers are all arguing about, in plain terms: how should a model deal with its own past? The original transformer’s answer — attention — is to re-read everything, every time, at full price. The alternative family this build belongs to keeps a state: a running summary, updated token by token — less like consulting an archive and more like continuously rewriting a notebook as you read. And because a summary must eventually let old things go, this family has one knob that matters more than the others: how fast to forget. The current argument between laboratories — Qwen’s version against Kimi’s — is about how many independent forgetting dials the machine should have: one per memory lane, or one for every slot inside each lane. Underneath the jargon (“decay gate”, “per-head”, “per-channel”), that is the entire dispute: how precisely should a machine forget?
The answer built here is a hybrid: 20 layers, of which 15 carry the recurrent state and 5 — every fourth layer — are full attention. The ratio is 3:1. The intuition: most of a sentence is local business, cheap to handle with a state; every so often, the model needs a long, direct look backwards, and attention provides it. 388 million parameters on the main path.
Every block of this model was cross-checked, layout by layout, against Qwen3.5’s published weights and reference implementation before any experiment touched it. The reason is not reverence. Innovation on an unverified base is indistinguishable from debugging — if anything can be wrong, a surprising result means nothing.
Day zero: a model that knows nothing, evenly
Training starts from billions of small, near-random numbers. And here is the first magic trick of the whole pipeline: you can measure total ignorance. Ask the untouched model to predict the next token and score its surprise in nats — the natural unit of “how wrong was that guess.” A model that knows nothing distributes its guesses evenly over the vocabulary, and even ignorance has an exact value: the logarithm of the vocabulary size. For 65,536 tokens, that is 11.09 nats.
The initialization landed on 11.09 to two decimals. That is the smoke alarm: any wiring bug, any mis-scaled layer, moves that number immediately. A model either knows nothing evenly, or it knows something — and on day zero, “something” always means a bug.
The next landmark is 7.76 nats — the score of a model that has learned nothing but the shape of the language: which letters exist, which tokens follow which, at the frequency real text uses them. It was crossed around step 200. Facts come later; the rhythm of the language comes first.
Anatomy of a training step
Everything after this is a loop. One step of training, decompressed:
And the loop repeats — two thousand times for one sweep of an experiment. The training log compresses each step into a single line:
Four numbers carry the whole health of the run. loss — average surprise; it must fall. lr — the learning rate, here still warming up toward its constant. gnorm — the size of the total correction; a spike means instability, a collapse means stuck, this one is calm. tok/s — throughput, the price of the experiment in hours: at ~10,000 tokens per second, one 0.52-billion-token arm takes about 15 hours on this single card. Frontier models run the identical loop with a budget a thousand times larger.
How do you know it’s working?
You cannot watch 388 million numbers change. You watch one: the held-out loss — the exam, run on frozen text that was reserved from the start, never trained on, checksum-verified to have survived every corpus edit intact. When the exam improves, the model generalizes; when only the training loss improves, the model is memorizing.
Two disciplines make that one number trustworthy. First, the noise band: repeated evaluations wander by roughly ±0.01 nat, so an effect smaller than that does not exist until two independent runs agree it does. Second, the paired checkpoint: versions are compared point by point, every hundred steps, same exam, same everything.
Here is what a real result looks like under those rules. Version 3 added a sparse memory — a 134-million-parameter hash table addressed by local 2- and 3-grams, the “have I seen this exact little pattern before” reflex, injected into the network before layer 2:
That last line is the one a scientist reads first. Needle retrieval improved sharply, token copying did not move at all — so the gain is specific, not a generally luckier run. A result you can attribute is the only kind worth having.
What iteration actually looks like
Everything above describes the loop as it should go. Here is how it actually went, one evening in September, when version 4 — the channel-wise decay gate — was launched for the first time.
A quick memory check had just reported a comfortable 20.9 GB. The real training run crashed within minutes: out of memory, at 22.9 GB. Halving the batch size crashed too, at 21.5 GB — barely lower. The cause was structural, and it is the kind of thing no paper mentions: with a per-channel decay, the tensor of decay ratios — (batch, heads, chunk, chunk, channels) — must be kept in memory for the backward pass. It is a factor of ~128 fatter than the per-head version, and it does not shrink proportionally with batch size. The comfortable 20.9 GB figure, it turned out, had been measured on the wrong model entirely — the launcher’s default architecture, not the candidate. A micro-benchmark is only evidence about the model you actually loaded.
The fix: don’t save the decay chain at all — recompute it during the backward pass, a technique called gradient checkpointing. And before the fix was trusted, it was verified bit for bit: the checkpointed and un-checkpointed model must produce byte-identical outputs and gradients. They did — diff 0.0. The run restarted at 20.8 GB and has been training since.
Version 4’s gate is structurally new — a low-rank projection producing 1,536 decay values per layer where the old one produced 12. The identity test answers the confound question: copy the old gate’s weights into the new machinery, and the new model must reproduce the old one exactly. It does, to 2.38×10−7. From that point on, whatever the sweep measures can only come from the gate — if the result is real, it is attributable; if it is noise, that is measurable too.
What this article left out
A base model — even a good one — is not yet the thing you talk to. Instruction tuning, preference training, tool use, safety behavior, and then the engineering of serving: quantization, caching, latency. That pipeline is real work, and with this model it is also genuinely still ahead: the current milestone is a base model mid-experiment, with a sweep running as this is written.
Which is the last thing worth saying about scale. None of what you just read required a datacenter. One consumer GPU, one careful method, and the discipline to write the verdict down before knowing it. The experiments continue: right now, the per-channel decay gate is being tested against the per-head one — four training runs, two seeds each, about four days of compute — and the reading criteria were fixed in writing before the first token. The next article in this series is that verdict.
It is built in two thousand careful steps, judged by a number you decided to trust in advance.
Written during an active training sweep on the author’s machine (a single RTX 3090), in collaboration with Kimi (Moonshot AI). Every number quoted is taken from the project’s logs and documents; nothing is reconstructed from memory. Pedagogical simplifications are deliberate; the events are not.
THE SERIES
- The 388-Million-Parameter Laboratory — the laboratory itself: four versions, one change each, and the rules for reading results.
- Next: the v4 verdict — channel-wise decay vs. scalar, and the 2×2 against the n-gram memory.
Prev: The 388-Million-Parameter Laboratory · Next: The Architecture Atlas · ← back to blog