Large language models (LLMs) exhibit strong general capabilities, yet their parameters typically remain fixed after deployment, limiting learning from new interactions. In open-ended environments, this motivates self-evolving agents that continually update reusable state-including model parameters, memories, tool defin...
Jia-Hao Chen, Zhou Feng, Ou-Bo Ma et al.· 0 citations
Quantile-Guided Density Estimation (QGDE), which approximates this distribution with multiple quantile trends and uses local density weighting to produce token-level estimates and suggests that released tokenizer vocabularies provide a useful signal for fine-grained corpus estimation beyond coarse composition inference...
Qingjie Zhang, Xing-Zhang Ren, Zi-Xuan Chen et al.· 0 citations
This work proposes Sampled-BPE, a lightweight token-level auditing pipeline that sample a small subset and train BPE tokenizer to surface polluted tokens, and releases a hierarchical Chinese web token dataset with 660k+ token records, organized as trees to support review and tracing of pollution.
Qingjie Zhang, Ziqi Tang, Jie Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.