This paper conducts a systematic study of unsupervised visual pretraining paradigms that directly leverage visual documents without text extraction, showing that Visual Pretraining is a scalable learner for foundation model intelligence.
Experiments on challenging reasoning benchmarks show that H$^2$SD achieves the strongest overall performance among representative RLVR and self-distillation baselines, with stable optimization and a favorable accuracy-efficiency trade-off.
Qi Cai, Yi-Chuan Ma, Linyang Li et al.· arXiv.org· 2 citations
Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings.
Lei Bai, Jiaqi Cao, Chiyu Chen et al.· 2 citations
This work presents Memory Decoder at Scale, scaling memory models up to 6.9B parameters and pretraining them on 300B tokens, demonstrating that independently scaling pretrained memory offers a more parameter efficient path to improving language model performance.
Rubin Wei, Jiaqi Cao, Jiarui Wang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.