Contract inference requires multiple judgments about a shared document, but aggregate accuracy can conceal changes in the individual decisions. Repeated agreement is also insufficient: a model may consistently return the wrong answer. In this paper, we compare Jev with nine language models on ContractNLI, evaluating in...
Fan Zhang, Yan-Kai Chen, Zhuo-Han Xie et al.· 0 citations
Automated prediction markets require sponsors to prefund liquidity before observing order flow, creating a financing challenge at launch. We study whether nonnegative charges conditioned on observable payoff direction can improve recovery of this prefunded capital while limiting their effect on informed participation....
Yankai Chen, Bowei He, Zhuohan Xie et al.· 1 citation
Generalizable dynamic graph anomaly detection (DGAD) enables pretrained detectors to identify anomalies in unseen target domains without costly retraining. However, existing methods often fail for two reasons. First, they mainly rely on domain-agnostic patterns and miss domain-specific patterns that keep evolving. Seco...
Jialun Zheng, Han-Chen Yang, Jiannong Cao et al.· 0 citations
This work introduces a unified conceptual framework that views discrete diffusion models through the construction of the underlying discrete state space, and exposes common design trade-offs across training objectives, inference algorithms, scaling behavior, systems optimization, and evaluation protocols.
Ye Yuan, Wei-En Li, Rui Song et al.· arXiv.org· 0 citations
Experiments on Chinese fund-market traces from 2021 to 2026 identify a stable leading group of LLM advisors that combines substantially stronger personalized content with competitive investor-side trajectory outcomes.
Jie Gong, Mao-Wei Jiang, Zhiwei Liu et al.· 3 citations· ⚡1
AdaMTP is proposed, an adaptive training paradigm that dynamically aligns the prediction horizon with the intrinsic predictability of the sequence, and consistently outperforms standard MTP in both task performance and inference speedup.
Ziqiang Cui, Han Shi, Bowei He et al.· 0 citations
A controlled harness evaluation of memory substrates for memory-augmented agents, covering dense and sparse indices, text records, structural stores, hierarchical stores, refinement-based memories, parametric updates, and activation-compatible context mechanisms, shows that no single substrate consistently dominates.
Wei-Chieh Huang, Wei-Zhi Zhang, Yu-Chen Wu et al.· 2 citations
BPO is instantiate as Branching Policy Optimization (BPO), a sandbox-native RL algorithm that adaptively snapshots the sandbox at high-entropy decision points along a backbone trajectory, and proves this estimator is unbiased and has strictly lower variance than the trajectory-level baseline, with the reduction equal t...
Bowei He, Yankai Chen, Xiaokun Zhang et al.· 1 citation
FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the correct answer to finance questions involving domain terminology, numerical interpretation, and conceptual financial reasoning across languages...
Zhuohan Xie, Yu-Yang Dai, R. Elbadry et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.