The Architecture Atlas
Contents
There is no single Transformer anymore
Ask someone what a modern language model looks like and the answer will usually be some variation of “a Transformer.” That answer is still useful, but it is no longer very precise. Since the original Transformer, researchers have spent years changing almost every part of the design: how tokens interact, how information is remembered, how much computation each token receives, where parameters live, and what gets recomputed at inference time.
Some of those changes are replacements. Many are not. Grouped-query attention does not replace the Transformer; it changes how its key and value heads are organized. Mixture-of-Experts does not require abandoning attention; it changes how parameters are activated. Sliding-window attention changes the connectivity pattern. Gated DeltaNet changes the memory mechanism itself. N-gram systems can even sit partly outside the neural network and still change the effective system we deploy.
The interesting architectural question is not simply “attention or no attention?” It is: what kind of memory, connectivity, and computation do we want each token to have access to? Modern architectures can be understood as different answers to that question.
The next architecture is often not a replacement for the previous one. It is the part that the previous one was bad at.
This makes the field easier to understand. Instead of memorizing dozens of names, we can follow the engineering pressures that produced them.
The Transformer: arbitrary retrieval as a superpower
The 2017 Transformer replaced recurrence and convolution in the core sequence model with attention. Its defining trick is simple to describe: for every token, construct queries, keys, and values, then use similarity between queries and keys to decide which values matter. This gives the model direct, content-dependent access to earlier positions. Attention Is All You Need ↗
This is the property that keeps attention so hard to beat: the model does not have to compress the entire past into one fixed state before consulting it. The past remains explicitly addressable. That flexibility is powerful, but it comes with a price. During training, full attention scales quadratically with sequence length, and during autoregressive inference the key-value cache grows with the number of stored positions.
Exact, content-based retrieval. A token can effectively ask “which previous information is relevant to me?” That is difficult for many linear-time or recurrent mechanisms to reproduce perfectly, especially when the answer depends on a specific piece of information buried far away in the context.
First pressure: make attention cheaper
The first response was not to throw attention away. It was to make attention less wasteful. Grouped-Query Attention (GQA) is the natural generalization of the standard design: several query heads share one key-value head per group instead of each having its own. The mechanism stays familiar, but the inference footprint becomes substantially smaller while quality stays close to conventional multi-head attention. GQA — Ainslie et al., 2023 ↗ Multi-Query Attention (MQA) is the extreme case of the same idea: every query head shares a single key-value head, which sharply reduces inference memory and bandwidth. One Write-Head is All You Need — Shazeer, 2019 ↗
This is an important pattern because it appears again and again: an architectural innovation does not need to change the conceptual role of a component. Sometimes it simply changes its cost.
There is an even more aggressive answer in the same direction. Multi-head Latent Attention (MLA), introduced with DeepSeek-V2, compresses the keys and values into a single low-rank latent vector that is cached at inference instead of the full per-head K/V tensors. Attention still happens — it is decompressed on the fly — but the cache shrinks well beyond what head grouping achieves. DeepSeek-V2 ↗
Second pressure: make the connectivity local
If every token does not need to inspect every other token, why pay for a global interaction pattern? Sliding Window Attention (SWA) limits each token's direct attention to a local neighborhood. Longformer is an early well-known example of sparse attention patterns that combine local windows with selected global positions. The trade is immediate: less computation and a smaller effective attention footprint, but less direct access to arbitrary distant tokens. Longformer ↗
The interesting insight is that “local” does not necessarily mean “short-context.” A model can propagate information across many layers. The architecture is changing how information travels, not simply deleting the past. Gemma 3 takes the local-global split literally: five sliding-window layers (window 1024, short-rotation positional encoding) followed by one global layer (long-rotation encoding) — a pure alternating hybrid with no recurrent state at all. Gemma 3 ↗
Learned sparsity: NSA and DSA
A sliding window is a hand-drawn connectivity pattern: the architect decides, in advance, that every token will see the same fixed neighborhood. Native Sparse Attention (NSA, DeepSeek) makes the pattern itself learnable. Three branches run in parallel over the key-value cache — a compression branch that pools distant KV into coarse blocks, a selection branch that keeps only the top-scoring blocks, and a local window branch — and a learned gate mixes their outputs per head. Trained natively from scratch rather than bolted on afterward, it won a best paper award at ACL 2025. NSA — best paper, ACL 2025 ↗
The idea shipped fast. DeepSeek-V3.2 (September 2025) replaced its attention with DSA, a descendant tuned for serving — one of the levers behind API prices dropping by up to three quarters — and GLM-5 (February 2026) combines MLA with DSA in a 744B-parameter model. Fixed windows tell every token the same story; learned sparsity lets the model decide, per head and per input, which connections are worth paying for. DeepSeek-V3.2 ↗ GLM-5 ↗
Third pressure: increase capacity without proportional compute
Mixture-of-Experts attacks a different bottleneck. Instead of forcing every token through the same feed-forward parameters, an MoE layer contains multiple experts and a router selects which ones participate for each token. A model can therefore contain a very large number of parameters while activating only a subset for any given token. Switch Transformer is one of the canonical demonstrations of this sparse-activation idea at scale. Switch Transformers ↗
The same logic can be pushed one level higher. Mixture-of-Depths routes individual tokens past entire layers, turning depth itself into a per-token compute budget rather than a fixed property of the architecture. We return to it in the closing section.
MoE is orthogonal to attention. “Transformer” and “MoE” are not opposing boxes. A model can use full attention, SWA, GQA, MoE, or several of these at once. One describes token connectivity; another describes parameter activation.
Once you see MoE as a parameter-allocation mechanism rather than a complete model architecture, many “hybrid” designs stop looking hybrid.
Fourth pressure: escape quadratic attention
The more ambitious family of approaches changes the memory mechanism itself. Linear attention, recurrent architectures, and state-space models try to represent the past in a compact state rather than retaining a full pairwise attention matrix. Mamba is a prominent example: a selective state-space model whose parameters depend on the input, allowing the state to selectively propagate or forget information while maintaining linear scaling with sequence length. Mamba ↗
This changes the fundamental question. Full attention asks the model to keep a rich, addressable record of the sequence. A recurrent or linear mechanism asks it to maintain a compact state that is updated as the sequence advances. RWKV showed earlier that attention-like time-mixing can itself be reformulated as a cheap recurrence — a useful bridge in the lineage toward delta-rule models. RWKV ↗
Information remains explicitly represented in the sequence or cache, making arbitrary retrieval natural but expensive.
A fixed-size or structured state can scale much more gracefully with sequence length, but compression can discard details that later become useful.
Local or global attention can handle retrieval while recurrent state handles persistent sequence processing at lower cost.
Gated DeltaNet: memory that can be rewritten
Delta-rule approaches take the recurrent idea further. Rather than treating the state as something that merely evolves according to a fixed transition, the delta rule provides a targeted update mechanism: change the stored state in response to the current key, value, and prediction of what the state should contain. Gated Delta Networks combine two complementary ideas—gating for rapid forgetting and the delta rule for targeted state modification. Gated Delta Networks ↗
This distinction is subtle but important. A useful recurrent memory cannot merely “remember more.” It needs a way to decide what survives, what is overwritten, and what deserves a precise update. That is why gating and delta-style writes fit together so naturally.
Kimi Delta Attention: pushing the hybrid further
Kimi Linear takes this family of ideas into a larger hybrid architecture. Its Kimi Delta Attention (KDA) extends Gated DeltaNet with finer-grained gating, while the full model combines KDA with Multi-Head Latent Attention (MLA) in a 3:1 ratio rather than attempting to eliminate attention altogether. The reported design targets both memory efficiency and long-context decoding, illustrating a broader direction: a model can assign different layers to different kinds of memory rather than insisting on one universal mechanism. Kimi Linear ↗
The question is no longer “which mechanism wins?” It becomes “which mechanism is best for which layer, timescale, or operation?” Kimi Linear is a useful example of this design philosophy: preserve an attention mechanism where it is valuable, and use a linear recurrent mechanism elsewhere.
Qwen3-Next: the other hybrid lineage
Kimi Linear is not the only bet on gated delta memory at scale — and not even the only one from this family. Qwen3-Next (September 2025) built its hybrid from the other side of it: Gated DeltaNet layers handle the cheap, continuous sequence processing, and full Gated Attention layers are interleaved in a 3:1 ratio — twelve blocks of three delta layers plus one attention layer, each followed by a routed feed-forward. Qwen3-Next — Qwen Team, 2025 ↗
Around that trunk, nearly every mechanism in this article shows up at once: an ultra-sparse MoE — 512 routed experts plus one shared, ten activated per token — keeps active compute near 3B parameters out of 80B, and a multi-token prediction module is trained natively: the same MTP trick described in section 12, now part of a shipped architecture. In February 2026, Qwen3.5 promoted this exact layout to flagship status — 397B total, 17B active, 256K tokens of native context extensible to 1M. Qwen3.5-397B-A17B ↗
The pattern is robust enough to survive ingredient swaps. MiniMax-01 replaced the delta layers with a linear Lightning Attention — seven linear layers for every softmax one — and Ling 2.5 keeps one heavy MLA layer for every seven light linear-attention layers. Same hybrid contract, different internal machinery. MiniMax-01 ↗ Ling 2.5 / 2.6 ↗
KDA refines Gated DeltaNet's per-head scalar decay gate into per-channel gating; Qwen keeps the simpler per-head gate but pairs it with full attention where Kimi uses MLA. Both independently land on a 3:1 recurrent-to-attention ratio. Different dials, same answer to the same question.
N-grams: the strange return of explicit local memory
N-gram methods look almost prehistoric next to modern neural architectures, and that is precisely why they are interesting. An n-gram is simply a sequence of n tokens. Instead of learning every local continuation only through neural weights, a system can exploit repeated token patterns directly. Modern inference stacks already use n-gram caches and matching strategies for speculative decoding, where an inexpensive mechanism proposes likely continuations and the main model verifies them. llama.cpp speculative decoding docs ↗
But n-grams are no longer confined to the edge of the system. In 2026 they moved inside the model, as a first-class memory primitive. Engram (DeepSeek, 2026) is the clearest statement of the idea — and a direct answer to the capacity question of section 06. Repeated token patterns are stored in explicit lookup tables: hashed n-gram keys map to learned vectors, and the retrieved vectors are injected into the residual stream at fixed layers, gated by the current hidden state. Engram — DeepSeek, 2026 ↗ The table grows the model's total parameters, but a token pays only for a few constant-time lookups — capacity without proportional compute. The production proof has already landed: DeepSeek-V4.1-Flash (September 2026) ships 196 billion Engram parameters — n-gram orders {2,3,4}, eight hash heads, host-side prefetch over RDMA — inside a 552B-parameter model with a one-million-token context. DeepSeek-V4.1-Flash ↗
A wave of 2026 work explores the corners of the same space: LongCat-Flash-Lite folds hashed n-gram embeddings into the input, LongCat ↗ NGM adds untrained n-gram memory on top of an existing checkpoint, NGM ↗ TN-gram tensorizes the tables to remove hash collisions, TN-gram ↗ and SDM sparsifies a gated delta state itself. SDM ↗ The structural point for this atlas is that the third kind of memory — explicit, cheap, pattern-level — is no longer a proposer standing outside the model. It is a component inside the trunk, spending capacity the way MoE does: more total parameters, almost no extra compute per token.
This is not the same job as DeltaNet or attention. An n-gram mechanism is excellent when the relevant information is a literal or near-literal local pattern. It does not understand the pattern in the way a neural representation does. But it can be extremely cheap, and in workloads with repetition that can be enough to matter.
Seen this way, n-grams stop being an awkward historical footnote. They become another point in the architectural design space: explicit, cheap, pattern-level memory.
Multi-Token Prediction: more signal per position
Every pressure so far lives inside the model: how tokens interact, what is remembered, which parameters fire. One more axis sits at the output — how many future tokens the model is asked to predict from a single position. The classic setup asks for exactly one. Multi-token prediction (MTP) adds auxiliary heads that predict the token after next, and sometimes the one after that, all sharing the same trunk. Better & Faster LLMs via Multi-token Prediction ↗
This buys two things at once. During training, each position supervises several future tokens, so the trunk receives a denser learning signal — representations that must anticipate more than one step tend to capture structure earlier. During inference, the auxiliary heads can act as a learned speculative proposer: they guess upcoming tokens, and the main model verifies them. DeepSeek-V3 explicitly keeps its MTP module in the released model for exactly this purpose. DeepSeek-V3 Technical Report ↗
N-gram proposers are table-driven; MTP proposers are learned. Both propose candidate continuations that the main model then verifies. And like n-grams, MTP does not change the trunk's memory or connectivity — it changes the training signal and the output contract. EAGLE completes the proposer family with a small learned draft head that predicts future features rather than tokens — verified by the main model in the same accept-or-reject way. EAGLE ↗
Recurrent depth: compute that loops
Every mechanism so far spends each budget once: a layer runs, a memory is written, an expert fires. One budget remains almost untouched — compute itself. The Universal Transformer proposed weight-shared depth recurrence: one shared block, applied repeatedly, refines every token in place, with an adaptive halting budget per position. Universal Transformer ↗ Follow-up theory showed the point is not parameter saving but computation: constant-size looped blocks can execute iterative algorithms and in-context learning procedures that a fixed stack of the same size provably cannot. Looped Transformers as Programmable Computers ↗ In 2025 the idea was scaled into a real language model: Huginn, a 3.5B-parameter looped-depth model trained on 800B tokens, where the number of loop iterations becomes a dial for test-time compute. Huginn ↗
The Recurrent Looped Transformer (RLT), a September 2026 tech report, pushes the loop in a new direction: across tokens. The decoder keeps a hidden state that crosses the prompt–response boundary. For each token, the previous final decoder output is merged — through a learned gate — with the token's fresh encoder representation, and the merged result passes through the decoder blocks. Three memories coexist: an encoder-derived global KV read by cross-attention, a per-layer sliding-window cache, and the recurrent state itself. RLT report ↗
The striking consequence is depth: after t tokens, the recurrent path has traversed t × L blocks while the number of blocks evaluated per token stays fixed. Recurrence buys effective depth with time instead of parameters. The evidence is early but instructive: at 26–29M parameters on algorithmic tasks, RLT generalizes far beyond training lengths where matched transformers collapse — 100% parity at 256 bits, state tracking at four times the training length — yet no language-modeling results, throughput numbers, or large-scale runs exist yet. On the map of this article it is a hybrid of explicit local memory (SWA), explicit global memory (cross-attention KV), and implicit recurrent memory — on the compute axis that mixture-of-depths approaches from an entirely different direction by routing tokens past layers. Mixture-of-Depths ↗
Most innovations are compositions
This is where the taxonomy gets much more interesting. Modern architectures are increasingly composites of mechanisms that solve different bottlenecks.
An architecture is a composition of choices — not a product name.
A hypothetical modern model could therefore be “a Transformer” and “an MoE” and “GQA” and “SWA” all at once. Another could combine Gated DeltaNet with occasional attention, a routed feed-forward block, and an auxiliary n-gram system. There is no contradiction because these mechanisms answer different questions.
The architecture matrix
One useful way to compare these families is to stop asking which one is “best” and instead ask where each one spends its budget.
How much of the past is retained, and in what form? KV cache, local context, recurrent state, latent representation, or explicit pattern tables.
Can a token see everything, a window, a sparse subset, or only information propagated through a state?
Is computation dense, conditionally routed, recurrent, memory-bound, or shifted away from the main accelerator?
How many parameters can the model hold, and how many does a token actually touch? Dense weights, routed experts, or lookup tables that grow capacity without growing compute.
Every architecture is effectively choosing where to place complexity. Attention stores more explicit information and spends more compute retrieving it. Recurrent systems compress the past and spend less on sequence length. MoE spends parameters selectively. Learned sparse attention spends connectivity selectively, letting training decide where to cut. N-gram systems spend almost nothing on representation and a lot on exploiting repetition. Multi-token prediction spends a few extra output heads to buy a denser training signal. None of these choices is universally correct.
Why attention still matters
It is tempting to read the history of efficient architectures as a story in which attention keeps getting replaced. The actual story is less dramatic. Attention remains extremely attractive precisely because its memory is explicit and its retrieval is content-dependent.
Even recent linear-attention work frequently reintroduces attention somewhere in the stack. Griffin combines gated linear recurrence with local attention. Griffin ↗ The Gated DeltaNet paper evaluates its own hybrids along the same lines — Gated DeltaNet layers with sliding-window attention (H1), or with Mamba2 stacked on top (H2). Kimi Linear combines KDA with MLA, and Qwen3-Next interleaves Gated Attention in a 3:1 ratio. These designs do not look like a failed attempt to escape attention. They look like attempts to reserve attention for the places where it earns its cost. Jamba pushed the same philosophy into a released large-scale hybrid, interleaving Mamba layers with attention and using MoE feed-forward blocks throughout the stack. Jamba ↗ NVIDIA's Nemotron 3 Super folds the vocabulary of this article into one open model: Mamba-2 layers, attention, latent MoE, and a shared-weight MTP head for speculative decoding. Nemotron 3 Super ↗ YOCO attacks the cache from yet another angle: a stack of local layers writes a single shared global KV cache that every later layer reads — the cache is written once, not once per layer. YOCO ↗ The contract is no longer hypothetical: DeepSeek-V4.1-Flash (September 2026) builds its causal encoder-decoder on it — twenty causal encoder layers, with the upper layers' key-value state projected from the encoder's final state, so the global cache is written once instead of once per layer. DeepSeek-V4.1-Flash ↗
The future may not be “attention versus recurrence.” It may be learning where recurrence is enough, where attention is necessary, and how the two should communicate.
A timeline of the design pressure
Replace recurrence and convolution in the core sequence model with attention, making arbitrary content-based retrieval highly parallelizable. Paper ↗
Reduce the cost of connectivity and KV storage through local attention, MQA, and GQA. GQA ↗
Increase total parameter capacity while routing only selected experts for each token. Switch Transformer ↗
Selective state-space and delta-rule models push toward linear-time sequence processing, while hybrid designs recover some of attention's strengths. Mamba ↗
Auxiliary heads predict several future tokens from one position, densifying the training signal and enabling learned speculative decoding. DeepSeek-V3 ↗
NSA's three-branch design makes connectivity itself learnable (best paper, ACL 2025); within months, DSA ships the idea inside DeepSeek-V3.2. NSA ↗
KDA, Kimi Linear, Qwen3-Next and Qwen3.5, sparse attention, and other designs increasingly treat attention, recurrent state, and sparse or routed memory as complementary tools rather than mutually exclusive architectures. Kimi Linear ↗
N-gram lookup tables move inside the model as a first-class memory primitive, spending parameters rather than compute on literal repetition. Engram ↗
and more like an argument about where memory, computation, and connectivity should live.
That is the lens I find most useful when looking at a new model. Instead of asking “what architecture is this?”, ask what the designers decided to preserve, what they decided to compress, what they decided to make sparse, and what they decided was worth paying for. Once you ask those questions, names like Transformer, MoE, SWA, Gated DeltaNet, KDA, n-grams, multi-token prediction, and learned sparse attention stop being isolated inventions. They become pieces of the same design space.
And that design space is still moving. The most interesting models are increasingly not trying to invent one perfect memory mechanism. They are trying to build a stack in which different memories handle different timescales, different retrieval problems, and different compute budgets. Even the compute axis itself is starting to move: Mixture-of-Depths routes individual tokens past entire layers, turning depth into a per-token budget rather than a fixed property of the architecture. Mixture-of-Depths ↗
REFERENCES
- Attention Is All You Need — Vaswani et al., 2017. The original Transformer.
- Longformer: The Long-Document Transformer — Beltagy et al., 2020. Local sliding windows plus selected global positions.
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Ainslie et al., 2023. Query/key-value sharing and grouped attention.
- Fast Transformer Decoding: One Write-Head is All You Need — Shazeer, 2019. Multi-query attention: one shared KV head for every query head.
- DeepSeek-V2 — DeepSeek-AI, 2024. Multi-head latent attention: the KV cache compressed into a low-rank latent vector.
- Switch Transformers — Fedus, Zoph, Shazeer, 2021. Sparse expert routing at scale.
- Better & Faster Large Language Models via Multi-token Prediction — Gloeckle et al., 2024. Auxiliary future-token heads: denser training signal, self-speculative decoding.
- DeepSeek-V3 Technical Report — DeepSeek-AI, 2024. MTP modules trained for signal, kept for speculative decoding.
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Gu & Dao, 2023. Selective state-space sequence modeling.
- RWKV: Reinventing RNNs for the Transformer Era — Peng et al., 2023. Attention-like time-mixing reformulated as a cheap recurrence.
- Griffin: Mixing Gated Linear Recurrences with Local Attention — De et al., 2024. Hybrid recurrence + local attention.
- Jamba: A Hybrid Transformer-Mamba Language Model — Lieber et al., 2024. Mamba + attention + MoE at production scale.
- Gated Delta Networks: Improving Mamba2 with Delta Rule — Yang et al., 2024. Gating + targeted state updates.
- Kimi Linear: An Expressive, Efficient Attention Architecture — Kimi Team, 2025. KDA + MLA hybrid architecture.
- llama.cpp speculative decoding documentation — practical n-gram cache and matching mechanisms.
- EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty — Li et al., 2024. A learned draft head over features; the third family of speculative proposers.
- Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language Models — Raposo et al., 2024. Per-token routing past entire layers.
- YOCO: You Only Cache Once — Microsoft, 2024. One shared global KV cache, written once, read by every later layer.
- MiniMax-01: Scaling Foundation Models with Lightning Attention — MiniMax, 2025. Linear-attention hybrid in a 7:1 ratio, 4M-token context.
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention — DeepSeek, 2025. Learnable three-branch sparsity; best paper, ACL 2025.
- Gemma 3 Technical Report — Google, 2025. A 5:1 local-global sliding-window hybrid with per-layer rotary frequencies.
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models — DeepSeek-AI, 2025. DSA sparse attention in production serving.
- Qwen3-Next-80B-A3B — Qwen Team, 2025. Gated DeltaNet + Gated Attention in a 3:1 ratio, 512-expert MoE, native MTP.
- GLM-5 — Zhipu AI, 2026. MLA combined with DSA sparse attention at 744B parameters.
- Qwen3.5-397B-A17B — Qwen Team, 2026. The Qwen3-Next layout promoted to flagship scale.
- Ling and Ring 2.6 Technical Report — Ant Group, 2026. The Ling-2.5/2.6 base: Lightning Attention + MLA in a 7:1 hybrid ratio, retrofitted from the Ling-2.0 1T checkpoint.
- Nemotron 3 Super — NVIDIA, 2026. Mamba-2 + attention + latent MoE + shared-weight MTP in one open model.
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models (Engram) — Cheng et al., DeepSeek, 2026. Hashed n-gram tables as conditional memory injected at fixed layers; the reference design for n-gram architectures.
- Scaling Embeddings Outperforms Scaling Experts in Language Models (LongCat-Flash-Lite) — Liu et al., Meituan, 2026. Input-side hashed n-gram embeddings as up to 46% of total parameters.
- NGM: A Plug-and-Play Training-Free Memory Module for LLMs — Qu et al., Nanjing University, 2026. Untrained n-gram memory on top of an existing checkpoint.
- Tensorizing Engram: Sharing Latents Across N-Gram Embeddings is Beneficial in LLMs (TN-gram) — Zhou et al., 2026. CP tensorization of n-gram tables, zero hash collisions.
- Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity (SDM) — Cabannes et al., Meta FAIR, 2026. Sparse reads and writes over a large explicit state, extending Gated DeltaNet.
- DeepSeek-V4.1-Flash — DeepSeek-AI, 2026. 552B MoE + 196B Engram; causal encoder-decoder after YOCO; n-gram orders {2,3,4} in production.
- Universal Transformers — Dehghani et al., 2018. Weight-shared depth recurrence: one shared block refines every token, with adaptive halting.
- Looped Transformers as Programmable Computers — Giannou et al., ICML 2023. Constant-size looped blocks execute iterative algorithms that fixed-depth stacks of the same size cannot.
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach (Huginn) — Geiping et al., 2025. Looped-depth pretraining at 3.5B parameters and 800B tokens; loop count becomes an inference-time compute dial.
- Recurrent Looped Transformer — Zhang, Feng & Qin, 2026. The loop crosses tokens: decoder state feeds back across the prompt–response boundary.
Previous: Behind the Curtain · ← back to blog