Can AI Agents be benchmarked within strictly controlled simulated trading environments to investigate how external factors impact their collective trading behaviors? These factors, which frequently influence trading behavior, are critical elements in the quest to maximize investors’ profits. Our work aims to address this problem by utilizing large language model-based agents. We have developed a multi-agent AI system, StockAgent, driven by LLMs and designed to systematically explore simulated trading behaviors in controlled environments. StockAgent enables examination of how external factors might affect agent behavior and profitability in simulations, without empirical validation for real-world use. Additionally, StockAgent avoids the test-set leakage issue present in existing AI-agent-based trading simulation systems. Specifically, it prevents the model from leveraging prior knowledge it may have acquired related to the test data. We evaluate different LLMs within the StockAgent framework, which serves as a benchmark for testing LLM behavioral tendencies with rigorous internal validity checks and non-LLM baselines. The experimental results demonstrate the impact of key external factors on stock market trading, including trading behavior and the rules governing stock price fluctuations. This research examines the phenomenon of agents’ free-trading gaps in the context of no prior knowledge of market data. The patterns identified through StockAgent simulations offer methodological insights into LLM behaviors in simulated financial environments. The code is available at: https://github.com/MingyuJ666/Stockagent.
Chong Zhang, Xinyi Liu, Zhongmou Zhang et al.· ACM Transactions on Intellig...· 0 citations
The workshop brings together researchers and practitioners from data mining, LLMs, NLP, NLP, IR, human-centered AI, and AI safety to position personalization as a central research direction for next-generation AI systems at KDD.
Xiaoyan Zhao, Yang Zhang, Moxin Li et al.· Proceedings of the 32nd ACM...· 0 citations
The field of information retrieval has been rapidly transformed by AI technologies, especially large language model (LLM) agents with strong reasoning, planning, and conversational capabilities. These AI agents have improved how information is retrieved, processed, and personalized across search and recommendation systems. Despite these advances, important challenges remain, including relevance, bias mitigation, real-time response, and data security. This workshop aims to bring together researchers and practitioners to discuss recent advances, practical applications, and future directions of AI agents in information retrieval, while encouraging collaboration and knowledge exchange within the community.
Qingsong Wen, P. Mehrotra, Yongfeng Zhang et al.· Proceedings of the 32nd ACM...· 0 citations
In vibe coding, people describe software in natural language and delegate implementation to AI agents. By analogy, vibe commerce allows people to express buying or selling goals in natural language and delegate the corresponding tasks to agents. Commerce, however, requires independently controlled Buyer and Merchant agents to interact in a shared market while preserving their private objectives and distinct authority. We introduce Agentic Commerce World (ACWorld), an environment for evaluating such agents across ongoing transactions. Through its Vibe Commerce Protocol (VCP), ACWorld validates agent actions before updating shared transaction state and records the resulting interactions, making agent behavior auditable and evaluation reproducible. The ACWorld Benchmark contains a 200-task capability-coverage track and a 60-task large-catalog track that searches 785,022 transactable listings. Across ten models, mean scores range from 65.9% to 85.6% and from 56.1% to 91.4%, respectively. Our analysis shows that process-level evidence is necessary: final state alone can miss evaluated errors, incomplete trajectories still retain useful process signals, and large-catalog tasks expose bottlenecks across stages.
Shichen Fan, Mingdai Yang, Duo Wang et al.· 0 citations
Large language models (LLMs) and agentic AI systems are rapidly moving into user-facing applications, yet most remain fundamentally generic, optimized for population-level objectives under the assumption that one model can serve all users. This assumption is increasingly misaligned with real-world deployment, where AI systems interact continuously with individuals whose preferences, knowledge, goals, and values evolve over time. PILA'26 is motivated by the need to move beyond static general models toward personal intelligence ---AI systems that explicitly model users and dynamically adapt their reasoning, behavior, and decisions through memory, interaction, and lifelong learning. The workshop brings together researchers and practitioners from data mining, LLMs, NLP, IR, human-centered AI, and AI safety to position personalization as a central research direction for next-generation AI systems at KDD. Topics include user memory and personalized alignment, self-evolving and lifelong learning, datasets and evaluation, real-world applications, and trustworthiness in user-adaptive AI. Workshop website: https://pila26-workshop.github.io.
Xiaoyan Zhao, Yang Zhang, Moxin Li et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.