Despite rapid progress in automating scientific research, generating promising and well grounded research solutions remains a central challenge. We isolate research ideation as a standalone task and build our solution on the intuition that a challenge in one field can often be addressed by a mechanism that solved an an...
Jia-Rui Liu, Ren-Jie Tao, Yi-Wei Liao et al.· 0 citations
We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends t...
Ru-Lin Shao, Shannon Zejiang Shen, J. Yin et al.· 0 citations
Switch Distillation is proposed, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy, which consistently outperforms existing distillation objectives across teacher sizes.
Jacqueline He, Howard Yen, S. Li et al.· 0 citations
Anchored Decoding is proposed, a plug-and-play inference-time method for suppressing verbatim copying that enables decoding from any risky LM trained on mixed-license data by keeping generation in bounded proximity to a permissively trained safe LM.
Jacqueline He, J. Hayase, Wen-tau Yih et al.· arXiv.org· 0 citations
This work introduces S-EMBER (Streaming Egocentric Memory Benchmark for Episodic Retrieval), a large-scale benchmark comprising 3,141 videos totaling 388 hours of organic activity captured via Ray-Ban Meta smart glasses that establishes a hardware-authentic foundation for developing grounded, reliable episodic memory i...
Xiaodong Wang, Xuanyi Zhao, Pedro Rodriguez et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.