Back to articles
AI Agents

AsyncLLM Turns LLM Inference into a General Asynchronous Agent

3 min read

Introduction

Many LLM agents still follow a simple cycle: receive an input, reason about it, produce a response or call a tool, and then wait for the next input. That pattern works well for chat and short automation workflows, but it becomes restrictive when the world keeps changing while the model is busy. A voice assistant may need to hear new speech while processing an earlier utterance. An embodied system may have to interpret its surroundings while planning a movement. A monitoring service cannot stop ingesting events simply because its previous analysis has not finished.

Yandex Research’s open-source AsyncLLM framework explores a more general way to build such systems. In the paper “LLMs are General Asynchronous Agents,” the authors argue that asynchronous behavior should not be limited to task-specific architectures or an orchestration layer around model APIs. Instead, concurrency can be made part of the inference process itself.

Key points

  • Concurrency is placed inside inference. AsyncLLM is not merely a wrapper that launches several independent API calls. It allows multiple inference processes to run at the same time and to read states that are still evolving.
  • Agents are expressed as coroutines. An agent can be described as a collection of Python async/await coroutines. Separate coroutines can observe inputs, analyze them, update memory, or perform actions, while standard events, locks, and queues coordinate their work.
  • Cache blocks provide shared state. Each coroutine writes to cache blocks and uses cache views to select which blocks it should attend to. This allows processes to share information without forcing every inference step to revisit the entire context.
  • The system supports different model components. AsyncLLM is built on a Mini-SGLang-based inference engine that batches requests across coroutines. The reported implementation supports full attention, Gated DeltaNet, and multimodal M-RoPE models.
  • No task-specific training is required in the demonstrations. The authors show Qwen 3.x models operating asynchronously for streaming video understanding, videogames, and monitoring. The central change is how inference is organized, rather than a separate training pipeline for each application.

Why it matters

In many existing agent systems, concurrency is handled by the application layer. The application starts several requests, while each model call remains essentially sequential. This can lead to repeated context transfer, state synchronization problems, and unnecessary latency. AsyncLLM treats concurrency as a native inference abstraction, allowing several lines of reasoning to work with overlapping, continuously updated memory.

That design is a natural fit for persistent video analysis, real-time monitoring, and interactive environments. One coroutine could keep consuming a visual stream, another could perform a slower and deeper assessment, and a third could respond to a detected event. These activities do not have to wait for one another at every step.

The approach also gives engineers a familiar programming model. Coordination is based on standard asynchronous primitives rather than a separate task-specific control language. At the same time, the framework does not by itself guarantee that every asynchronous workload will be faster or more reliable. Shared state can conflict, stale plans may need to be interrupted, and additional concurrent reasoning can increase compute usage. Deciding which process should access which memory remains a key systems problem.

AsyncLLM therefore represents an architectural direction rather than a finished solution to all real-time agent challenges. Its broader contribution is a shift in perspective: an LLM need not read everything, think once, and then act. It can instead participate in several coordinated processes that observe, reason, and respond as new information continues to arrive.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles