One Year of Online Retail: How E-Commerce Bench Tests Long-Horizon Agents
Introduction
A model that can search for a product, write an advertisement, or call a few tools has not necessarily demonstrated that it can run a business. E-commerce is difficult because decisions are connected over time. A purchasing choice affects inventory weeks later; a discount can create sales while putting pressure on cash; and a return or supply disruption may expose weaknesses long after the original decision. E-Commerce Bench is designed to measure that longer chain of consequences.
From task completion to year-long operation
The benchmark asks an LLM agent to operate multiple online stores over 365 days. The agent must research the market, choose and source products, negotiate with suppliers, adjust sales strategies, fulfill orders, process returns, and manage cash flow. Its final objective is to maximize total assets at the end of the year. This setup differs from short tool-use tests because early actions can shape later inventory, liquidity, and profitability.
The environment changes throughout the calendar. Product and supplier records are derived from a real e-commerce platform, while promotions, natural disasters, and supply-chain shocks alter demand over time. A successful agent therefore needs more than a fixed pricing rule. It must observe outcomes, learn from experience, and revise its operating policy as conditions change.
What the benchmark measures
- Multi-round negotiation: Procurement involves supplier pricing, concessions, and decisions rather than a single quote comparison.
- Long-term coherence: The 365-day horizon tests planning, state tracking, and consistency across many interactions.
- Multiple dimensions: Eighteen frontier models are evaluated across seven dimensions instead of being ranked only by final assets.
- Reproducibility: Customer purchases and returns follow a fixed demand model. A negotiation kernel determines supplier prices, concessions, and decisions, while the LLM is used mainly to verbalize those outcomes.
The available abstract reports that there is no universal winner. GPT-5.6 Sol grows an initial stake of 100,000 to 1,431,425, the highest ending result reported, but ranks 16th out of 18 on fraud avoidance and trails Fable5 i… in that dimension. Because the supplied abstract is truncated, the remaining comparison should be checked against the full paper rather than inferred here.
Why it matters
The benchmark moves agent evaluation toward sustained business operation. It makes clear that an autonomous commerce system should not be judged by revenue or final assets alone. Risk avoidance, compliance-related behavior, inventory discipline, liquidity management, and adaptation to shocks are also part of operational competence.
The deterministic market is a strength for controlled comparison, but it also limits how directly the results transfer to live commerce. The benchmark primarily isolates planning, memory, tool use, and policy adaptation under repeatable conditions. Future evaluations could combine this reproducibility with more open competition and publish full trajectories across each business dimension. For companies, the practical lesson is straightforward: the agent that earns the most may not be the agent that is safest to entrust with an actual store.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...