LLM agents increasingly maintain personal memory across sessions, but it can conflict. Preferences depend on context, behavior evolves, and sources can conflict. When a query lacks context, time, or source authority to interpret conflict, treating one memory as definitive converts unresolved conflict into an unjustifie...
Lulu Yang, Shusheng Xu, Zhuo-Ran Li et al.· 1 citation
Reinforcement learning (RL) post-training often uses distinct GPU kernels for rollout and policy update. In synchronous PPO and GRPO, numerical disagreement can perturb ratios between current token probabilities and those assigned during rollout. Recomputing rollout log-probabilities with the policy-update backend avoi...
Ran Yan, You-He Jiang, Jia-Yi Nie et al.· 0 citations
AReaL-DTE is presented, a snapshot-free Delta Transfer Engine that translates inference-visible weight sparsity into end-to-end system efficiency and achieves speedups of up to 19.9x over ByteCheckpoint and 3.2x over PULSE across clusters, and up to 7.6x and 7.4x within a cluster.
Yingqi Peng, Jia-Wei Zhang, Wenhao Zhou et al.· 1 citation
DualDecoder is presented, a lightweight serving system for long-context LLM inference that enables efficient sparse KV cache retrieval from host memory that leverages a novel dual-token decoding pipeline that accurately identifies critical KV entries with negligible computational overhead.