Results show that long-horizon reflective data is an effective route toward self-improving agents, and synthesize long-horizon improvement trajectories from machine learning and algorithmic programming tasks, two domains that offer verifiable feedback and reward sustained iteration.
Hong-Jin Qian, Chao-Fan Li, Kun Luo et al.· 0 citations
We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coherent patch reconstruction and scale its training with carefully curated generation...
Hai-Wen Diao, Jia-Hao Wang, Chen-Jing Ding et al.· 0 citations
This work introduces AREX, a family of Recursively Self-Improving (RSI) deep research agents that substantially outperforms comparable-scale baselines and remains competitive with models using substantially more activated parameters.
Shuqi Lu, Chaofan Li, Kun Luo et al.· arXiv.org· 2 citations· ⚡1
This work introduces VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable, and identifies recurring failure modes of the prevalent VLM-as-a-judge paradigm.
Junhua Xu, Rui-Si Wang, Fanyi Pu et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.