Back to articles
Memory & Context

Periscope Lets Frozen Language Models Read Beyond Their Context Window

3 min read

The long-context bottleneck

A larger context window does not automatically make a language model better at reading long documents. Attention and key-value cache costs grow quickly, while accuracy can deteriorate as irrelevant material accumulates. A team from KAUST presents Periscope as a different way to handle the problem: instead of changing the model or training a long-context adapter, it uses a frozen model through a sequence of short inference probes.

Factorizing the read

Periscope is designed for decisions over a finite set of answers, such as choosing the relevant document, selecting a supported option, or identifying the passage that contains evidence. The input document is divided into N chunks and placed on an approximately K-by-K grid, with K close to the square root of N. The system then constructs two families of probes.

Local probes contain consecutive chunks, preserving nearby wording and details. Strided probes sample chunks across the document, giving each query a broad view of the overall text. Both probe types ask the model the same question. Rather than requiring a full generated response, the method compares the log-odds of the candidate answers at a single output token. Each answer receives its strongest local and strided scores.

The same calculations produce an evidence map. By attributing probe scores back to the chunks they contain, Periscope estimates which parts of the document support each answer. The highest-scoring chunk can then be read in detail, allowing a long-document workflow to retrieve a small set of likely evidence before spending more computation on generation or verification.

Results and memory savings

The reported experiments suggest that selective reading can compete with brute-force windowing. On LongBench v2, reading only the roughly 9k tokens ranked highest by the map matched the same model’s best window-based result across windows ranging from 32k to 1M tokens. On InfiniteBench, whose median context is about 150k tokens, Periscope exceeded the best window-read baseline by 5 points. On BRIGHT’s long-document collections, it achieved the best NDCG@10 among six evaluated methods.

The memory profile is particularly notable. Each call caches only one probe, rather than the entire document. The authors report that a 27B model can process a 4.5M-token context on a single 80GB GPU, whereas a single direct pass would require 296GB of cache. This does not make computation free: the system performs multiple probes, and the total cost grows with document size. It does, however, shift the hardware requirement from “a GPU large enough for the text” to “a GPU large enough for the model.”

Why it matters

Periscope points to a useful distinction in long-context inference. Many tasks do not require the model to represent every token equally; they require it to find the few regions that determine a finite-choice decision. Combining global coverage with local resolution offers a practical alternative to endlessly enlarging the context window.

The method also has boundaries. Its strongest formulation assumes that the task can be scored over a known set of candidate answers. It is therefore not a direct replacement for unrestricted long-form summarization or generation. Chunk size, probe design, and question format can affect both accuracy and cost. In practice, Periscope is best viewed as a document triage and evidence-localization layer that precedes close reading or answer generation.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
RealCompanion Tests Whether AI Can Understand People Over Time
Memory & Context
cctest.ai
Memory & Context

RealCompanion Tests Whether AI Can Understand People Over Time

RealCompanion introduces a longitudinal benchmark for AI companions, evaluating not only whether systems can retrieve past messages but whether they know when the past matters. Its results suggest that memory is needed less often than standard benchmarks imply—and is still difficult to trigger correctly.

Read more