What Production Data Reveals About LLM Trading Agents
Introduction
When large language models are connected to live markets, the prompt is only one part of the trading system. The interface that exposes risk, the way rankings are displayed, and the tools available at each step may shape behavior just as strongly. A new production study from DXRG AI follows two trading-agent fleets over roughly six months to examine what the agents actually did after deployment.
What was measured
The first system, DX Terminal Pro, contained 3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days. The second, the DXAP live alpha fleet, included 500 to 599 user-created agents over its full history, with 91 to 117 active concurrently, trading Hyperliquid perpetuals. Together, the record includes about 7.5 million single-model invocations, roughly 300,000 onchain actions, and 231,638 multi-tool turns that produced 14,596 fills.
Four findings stand out:
- The operating layer strongly affects behavior. Each step on a risk slider was associated with about 0.425 additional leverage. Agent fixed effects explained roughly 60% of the observed variation. A leaderboard display boundary also altered selection: the probability of choosing an agent changed by about 1.75 times around the top-three cutoff.
- Position sizing was largely blind to volatility. Median leverage remained 5.0x across every volatility sextile. One posture-slider cell represented 11% of the book but accounted for 62% of liquidations, concentrating risk in a narrow configuration.
- Agents failed to retain much of the upside they reached. At least 43.2% of positions experienced a favorable excursion of 300 basis points or more within 24 hours. Yet 49.3% of those positions eventually closed with a negative trade return. A mechanical bracket comparison recovered 39.0 basis points per position.
- There was no demonstrated directional edge. The DXAP fleet was unprofitable and posted a 41% round-trip win rate, below a matched Hyperliquid retail benchmark at 50%. In a paired replay of 416 captured production scenarios, frontier model families were statistically indistinguishable in decision quality over this horizon, although their choice stability differed substantially.
Why it matters
The paper shifts the evaluation target from model intelligence alone to the complete agent stack. A risk slider is not merely a user preference; it can systematically alter leverage. A leaderboard is not neutral presentation; it can route user selection. Exit logic, sizing constraints, and volatility-aware controls may therefore matter more than switching between similarly capable frontier models.
The evidence should still be interpreted carefully. Both fleets came from one design lineage, operated in different markets and windows, and the mechanical bracket result is a comparative rule-based estimate rather than a promise of live returns. The study does not prove that every LLM trading system lacks an edge. It does show that, under these production conditions, the main weaknesses were execution and risk-control failures rather than a simple shortage of model sophistication.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...