Building strong chart-to-code systems increasingly relies on reinforcement learning, whose effectiveness depends critically on the quality of the reward signal. Large Multimodal Models (LMMs) play a natural critical role in jointly assessing chart visual appearance and task requirements. They are therefore increasingly...
Li-Jian Wu, Henry Hengyuan Zhao, Zi-Jian Zhang et al.· 0 citations
Generating text-rich images from prompts requires both textual fidelity and the coherent integration of text into the surrounding image. An explicit layout can provide structured guidance about what text should appear and where, but a well-formed plan alone does not guarantee that the renderer will realize it faithfull...
Guan-Qiao Chen, Jing Tan, Dong-Xing Mao et al.· 0 citations
VectorHarness is presented, a multi-agent framework for raster-to-authoring reconstruction that recovers heterogeneous components using type-appropriate native representations that recovers heterogeneous components using type-appropriate native representations.
Jia-Hao Tang, Yi-Ren Song, Alex Jinpeng Wang· 0 citations
MolDBG achieves competitive performance across all three tasks while enabling site‐specific affinity prediction and interpretable binding‐site discovery, and generalizes to structurally elusive targets, including cryptic pockets and intrinsically disordered proteins.
Gang Luo, Qian-Qian Zhang, Chen-Hao Wang et al.· Advancement of science· 0 citations
DAFS (Dynamic Attention-based Budget-aware Frame Selection), a training-free frame selector that improves over uniform sampling by up to 6.4 points on Video-MME and outperforms prior training-based selectors under matched frame budgets, while generalizing across selector and answerer backbones, and across tasks, withou...
Yilin Wang, Xiangxi Zheng, Dongxing Mao et al.· arXiv.org· 2 citations
Evaluating representative proprietary and open-source multimodal models, it is found that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks.