Skip to content

Can Agentic Trading Systems Pay for Their Own Intelligence?

Jul 2026 · arXiv.org · Vol abs/2607.10286 · 0 citations · 73 references
Computer Science

TL;DR

TradeLens is introduced, a trace-grounded diagnostic toolkit for evaluating agentic trading systems from their trading records, runtime traces, and deployment configurations, which reframe the evaluation of LLM-based trading agents from capability-centric performance ranking to trace-grounded diagnosis of intelligence-to-profit conversion.

Abstract

Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce trading value. Existing evaluations typically report performance metrics, but rarely examine agentic viability: whether dynamic LLM-mediated decisions convert their induced costs into measurable incremental profit. To apply this criterion, we introduce TradeLens, a trace-grounded diagnostic toolkit for evaluating agentic trading systems from their trading records, runtime traces, and deployment configurations. It reconstructs trading trajectories, attributes profit and cost to interpretable evidence, and diagnoses whether and why an agent pays for its own intelligence. We conduct extensive analysis across backbone models, capital scales, trading frequencies, and system architectures, together with deployment discussion. Our results show that viability hinges on intelligence-to-profit conversion: models exhibit different failure patterns, such as poor asset selection in DeepSeek-V3.2 and negative timing in GLM-4.7, while capital scale, trading frequency, and architecture matter only by amplifying or degrading decision-attributed timing value. These findings reframe the evaluation of LLM-based trading agents from capability-centric performance ranking to trace-grounded diagnosis of intelligence-to-profit conversion. Our code is available at https://anonymous.4open.science/r/TradeLens.

View source

Similar papers

Review Aug 2026

Agentic Quantitative Trading: A Survey of Workflows, Systems, and Evaluation

Quantitative trading is moving from isolated predictive models toward agentic workflows that combine reasoning, tool use, memory, and feedback. This survey reviews agentic quantitative trading across five stages: factor mining, signal discovery, portfolio construction, order execution, and risk management. We further e...

Feng-Rui Hua, Heng-Yi Yang, Xinqing Hao et al. · 1 citation
Preprint Aug 2026

TradingMoE: Routing the Right Experts in Evolving Markets

Large language models (LLMs) have shown strong potential for financial analysis and trading, but direct trading remains challenging because the predictive capabilities required can vary across assets, decision fields, and market conditions. Existing LLM-based trading systems either coordinate human-defined external exp...

Chang Zhou, Xingtong Yu, Minbin Huang et al. · 0 citations
Sep 2026

When AI Meets Finance (StockAgent): A Benchmark for Simulating Large Language Model Behaviors in Controlled Trading Environments

Can AI Agents be benchmarked within strictly controlled simulated trading environments to investigate how external factors impact their collective trading behaviors? These factors, which frequently influence trading behavior, are critical elements in the quest to maximize investors’ profits. Our work aims to address th...

Chong Zhang, Xinyi Liu, Zhongmou Zhang et al. · 0 citations
Preprint Aug 2026

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

Business Arena, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon, is introduced, taking a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.

Yijun Pan, Yu-Kun Lian, Kun-Yu Shi et al. · 1 citation
#artificial intelligence Preprint Sep 2026

What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets

We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, February to March 2026) and the DXAP live alpha fl...

T. Barton, Chris Constantakis, Patti Hauseman et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Agentic Empirical Asset Pricing: Methodological Foundations

This work evaluates SEADS against five re-implemented baselines on two US equity panels using this standard: no single metric ranks the systems consistently, motivating evaluation on multiple axes at once.

Ying-Jian Pan, Xiao-Wei Ding, Kay Giesecke · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.