General-purpose agents can plan, reason, and act over long horizons, yet their production capabilities remain fragmented across text, images, audio, video, documents, 3D assets, and code. Extending a foundation model to additional modalities ties capability growth to costly model updates, while assembling specialist mo...
Yan-Lin Li, Ming-Yang Hao, Sheng-Qiong Wu et al.· 0 citations
While symbolic regression (SR) has been successfully used in science to discover new equations, its use in financial valuation is hindered by several limitations. Whereas the natural sciences provide objectively correct relationships, financial valuation constitutes a distinct class of symbolic discovery problems, as i...
Kelvin J. L. Koa, Filip Orestav, Sheng-Qiong Wu et al.· 0 citations
The results support on-policy self-distillation as an efficient and analyzable approach to diffusion post-training by converting image-level reward guidance into explicit and continually refreshed intermediate supervision, thereby opening a path toward more efficient and diagnosable alignment.
Weina Zhou, Xiongwei Zhu, Ling-Dong Kong et al.· 4 citations
This paper proposes UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data, and hopes it will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.
Tianjie Ju, Zheng Wu, Yueqing Sun et al.· 1 citation
V-RAE, a video representation autoencoder that builds compact generative latents on top of frozen vision foundation model representations, and tFVD, a temporal-coherence diagnostic that correlates more reliably with downstream generation quality are introduced.