LongCat-DeepResearch Turns Research into a Coordinated Multi-Agent Workflow
Introduction
The challenge in deep-research AI is not simply retrieving more sources. A useful system must decompose an open-ended question, gather relevant evidence, connect findings across subtopics, and turn them into a coherent report. LongCat-DeepResearch, described in a technical report from the LongCat team, approaches this challenge as a workflow-design problem. Instead of asking one model to search, reason, write, and revise in a single chain, it assigns different stages to cooperating agents.
How the workflow works
The system is organized around four main steps:
- Planning before investigation: Several planning agents explore external sources and develop complementary views of the question. Their results are consolidated into an actionable research plan called ResearchSpec.
- Parallel section research: Research agents receive separate assignments based on the plan. They investigate their topics, collect additional evidence, and draft sections in separate contexts, allowing several lines of inquiry to proceed at once.
- Global review after assembly: Once the sections are combined, the system reviews the report as a whole. This stage looks for problems in coverage, reasoning, and connections between sections.
- Targeted local revision: Instead of repeatedly rewriting the entire document, the workflow directs changes to the sections that need attention. The intended benefit is to preserve global coherence while reducing unnecessary full-report regeneration.
The workflow also has a training use. Research tasks and trajectories produced during the process can support the mid-training and post-training of LongCat’s general-purpose models. In this sense, the research system is not only an application layer; it can also help create training material for improving the underlying model.
Results and what they show
LongCat-DeepResearch reports scores of 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II, and 79.83 on ResearchRubrics. On an in-house benchmark, it scores 76.04 and ranks second among four compared systems. These results indicate that structured planning and division of labor can support strong research outputs, but they do not establish that the system is best for every task or deployment setting.
The development analyses are particularly informative. Combining multiple planning perspectives tends to help, suggesting that early task decomposition can influence the quality of later investigation. However, further planning refinement produces mixed effects. More planning is therefore not automatically better. Additional editing improves average automatic readability preference across two benchmarks, but the direction of change differs between them. This points to a continuing trade-off between readability, evidence coverage, and analytical quality.
Significance and open questions
The report’s main contribution is a shift from single-pass generation toward a more inspectable research process. Section-level collaboration is a practical fit for broad questions with distributed evidence, while local revision offers a way to control the cost of editing long reports. At the same time, multi-agent workflows introduce coordination overhead and new consistency risks. The supplied material does not provide full details on runtime cost, latency, the compared systems, or human verification of factual claims. Those factors will be important for judging the system’s practical advantage beyond benchmark scores.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...