Offline replay should estimate what a discovery system could retrieve at a historical point, yet freezing the corpus leaves interaction memory unconstrained. We formalize point-in-time (PIT) discovery through historical state $(D_t, \theta_t, M_{<i})$ and introduce a paired replay that changes only memory availability....
Yi-Xi Zhou, Fan Zhang, Si-Kun Wang et al.· 0 citations
Contract inference requires multiple judgments about a shared document, but aggregate accuracy can conceal changes in the individual decisions. Repeated agreement is also insufficient: a model may consistently return the wrong answer. In this paper, we compare Jev with nine language models on ContractNLI, evaluating in...
Fan Zhang, Yan-Kai Chen, Zhuo-Han Xie et al.· 0 citations
Automated prediction markets require sponsors to prefund liquidity before observing order flow, creating a financing challenge at launch. We study whether nonnegative charges conditioned on observable payoff direction can improve recovery of this prefunded capital while limiting their effect on informed participation....
Yankai Chen, Bowei He, Zhuohan Xie et al.· 1 citation
This work introduces FinCUABuildBench, a benchmark for evaluating financial CUA task construction, and introduces FinCUABuildAgent, a multi-agent system for automatically constructing dynamic financial CUA evaluation tasks.
Jing-Pu Yang, Feng-Xian Ji, Jinri Guo et al.· 0 citations
TradeLens is introduced, a trace-grounded diagnostic toolkit for evaluating agentic trading systems from their trading records, runtime traces, and deployment configurations, which reframe the evaluation of LLM-based trading agents from capability-centric performance ranking to trace-grounded diagnosis of intelligence-...
Qiqi Duan, Changlun Li, Chen Wang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.