Wall Street Tests Eight Major AI Agents: Why Qwen Office Ranked First
Introduction
The AI Agent race is gradually moving beyond a simple comparison of model intelligence. The more practical question is whether an agent can complete a complex task reliably, use external tools correctly, and deliver an artifact that people can actually use. A recent assessment by Jefferies analysts tested eight major global agents in office scenarios. Alibaba’s Qwen Office ranked first overall and also received the highest estimated score for the engineering layer surrounding its model.
The result should be read as an observation from a specific benchmark, not as a universal ranking for every enterprise workflow. Still, it offers a useful view of how agent products are being evaluated.
Five tasks that test the full workflow
The assessment included five office tasks: summarizing an annual report from multiple files; searching online and comparing company operating data; controlling a real desktop browser to find information and generate a document; creating an English presentation from data; and producing a marketing poster based on a reference image.
According to the report, Qwen Office performed particularly well on complex office work, browser control, and multimodal generation. It was also the only product to score above 90 in every reported evaluation dimension. Such tasks test much more than fluent text generation. An agent must understand materials from different sources, decide when to browse, interact with software, recover from mistakes, and produce a structured final deliverable.
The hidden layer: Harness quality
Jefferies separates an Agent’s capabilities into two parts: the underlying model and its Harness. The latter refers to the engineering and product system that organizes the model’s work. It can include task instructions, context management, tool calling, boundary and permission controls, feedback loops, error recovery, and governance.
The report estimates that Qwen Office had the highest “implicit Harness score” among the eight products, ahead of tools including Claude Cowork and Codex. This distinction matters because a powerful model does not automatically produce a reliable business workflow. Long-running tasks can fail because of poor task decomposition, excessive context, incorrect tool use, or a lack of recovery mechanisms. A well-designed Harness can turn general model intelligence into a more predictable result.
Cost per task becomes a commercial metric
Agents often require multiple reasoning steps and repeated tool calls. They may also run for longer periods while maintaining context. As a result, enterprises will likely need to measure the total cost of completing a real task, rather than focusing only on the price of a single model request.
The source material notes that the Qwen 3.8 Max API used by Qwen Office is priced below some leading overseas models. Jefferies therefore sees a potentially attractive combination of performance and cost. However, API pricing alone does not determine business value. Success rates, human review, permissions, integration work, and the cost of correcting errors must all be included in the calculation.
Enterprise competition will center on workflows
As agents move from personal productivity to enterprise deployment, their defensibility may come from workflow and ecosystem integration. Connectors, Skills, historical tasks, collaboration tools, and permission structures can accumulate inside a product, increasing user stickiness and switching costs.
Qwen Office has initially connected with DingTalk’s instant messaging capabilities, supporting functions such as group-chat summaries, document and spreadsheet creation, and message or email handling. The larger test will be whether it can connect to company databases and real business processes.
The broader lesson from this assessment is not simply that one product won a ranking. Model capability is only the starting point. Harness design determines execution reliability, while task economics and enterprise integration determine whether an Agent can scale in production.
Source: QbitAI
Comments
Checking sign-in status...
Loading comments...