Recurrent Looped Transformer Makes Computation Grow with Sequence Length
Introduction
Transformers process sequences in parallel, but the number of layers applied to each token is generally fixed. That can be limiting for tasks such as parity computation, state-machine simulation, and permutation tracking, where the model must update an internal state repeatedly as new inputs arrive. Longer sequences may therefore require a computation path that grows with the input rather than a fixed per-token depth.
The paper introduces the Recurrent Looped Transformer (RLT), a design that adds sequential state propagation without simply increasing the fixed depth assigned to every token.
How RLT works
RLT divides an eight-layer network into two parts. A parallel causal encoder first produces a representation using the current token and its history. A recurrent decoder then combines that representation with the final decoder state from the previous token.
This creates two forms of computation at once. The encoder retains the parallel processing advantages of a causal Transformer, while the decoder passes a learned state forward one token at a time. As a result, the effective computation path becomes longer as the sequence grows, while the local cost per token remains comparatively fixed.
The researchers also vary how the eight layers are split between the encoder and decoder. This makes it possible to study the balance between broad parallel representations and the recurrent capacity needed for state tracking.
Main findings
The study evaluates five layer splits across six algorithmic tasks and compares them with an eight-layer standard Transformer. Results are reported over three random seeds.
- Parity: After training on sequences of at most 40 bits, two RLT configurations generalize to 256 bits with 100% accuracy for every seed. The standard Transformer remains near chance.
- S_5 permutation tracking: At eight times the training length, RLT reaches 97% final-state accuracy, while the standard Transformer stays below 1%. Performance also improves as the decoder receives more layers.
- Modular arithmetic: Beyond the training lengths, the best RLT configuration reaches 93%, compared with 33% for the standard Transformer.
- Feedback ablation: Removing recurrent feedback reduces both parity and swap-based S_5 performance to chance across all layer splits. The feedback is therefore central to the observed generalization.
Parallelism versus precise state updates
The authors also test a chunked variant that updates feedback once every four tokens. This allows tokens within a known chunk to run in parallel. On 64-bit parity, the approach retains 99% accuracy, suggesting that some tasks can tolerate less frequent state updates.
Permutation tracking is less forgiving. On length-64 swap-based S_5, chunking reduces accuracy from 100% to 20%. The contrast indicates that the right feedback frequency depends on the algorithm: aggregate properties such as parity may survive delayed updates, while order-sensitive state transitions require token-level recurrence.
Why it matters
RLT offers a useful alternative to simply adding more Transformer layers. It turns part of depth into recurrent time, allowing the computation path to expand with the sequence. The results support the idea that extrapolation is influenced not only by model size, but also by whether the architecture can maintain and update state at the required granularity.
The evidence is still limited to algorithmic benchmarks, so it should not be treated as proof of improved general language modeling. The trade-off between parallel execution and precise feedback also remains an important systems-design question.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...