In Hybrid Language Models, Attention Recalls and Recurrence Shapes Expression
A study of Qwen3.5 and Falcon-H1 finds a sharp functional split between the KV cache and the fixed-size recurrent state. Attention retrieves what appeared in context, while recurrence largely determines the language and persona used in the response.
Read more