Fisher-IRG yields stronger semantic-versus-nuisance predictive selectivity and generally more reproducible subspaces than covariance-based geometry, while recovering systematically distinct local directions.
This work contrasts a neutral question, a soft commercial instruction, and an explicitly adversarial instruction to ask about the sponsor's advantage while omitting the rival's advantage to establish neither typical behavior under advertising incentives nor effects on actual consumers.
This work introduces Video2Skill, a benchmark that covers robot tabletop manipulation and human kitchen activity and tests three core capabilities: locating manipulation events in time, grouping events of the same transformation, and deciding when to reuse an existing skill or create a new one.
Jian-Shu Zhang, Ce Zhang, Xi-Yuan Yang et al.· 0 citations
A holistic framework based on learnable information gain, which measures how much novel, parameterizable information a round provides relative to the previous round, and proposes ATRI (Adaptive Training Regulation via Information-gain), which reweights samples within a round and halts training across rounds when inform...
Chen-Xu Wang, Chao-Zhuo Li, Xin-Ze Shi et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Across six diffusion models, linear probes distinguish passing from failing attempts, with the strongest reads generally appearing beyond the early layers, and probe point estimates offer no consistent advantage in comparisons with model confidence.
A. Miglani, Samrath Singh Chadha, Kevin Li et al.· 0 citations
A unified theory for RE(S) is developed that covers the full spectrum of S, and can be interpreted as a stage-wise optimization process, where each stage takes $S$ gradient steps for minimizing the Kullback-Leibler distance to a fixed reward-weighted rollout distribution.
Zhi-Wei Wang, Yan-Xi Chen, Ya-Liang Li et al.· 0 citations
CaRE-KD is proposed, a confidence-gated distillation framework that replaces static objectives with uncertainty-adaptive optimization and provides a gradient-level analysis showing how this dual-granularity design induces a conditional calibration mechanism that prior static divergences cannot reproduce.
Experimental results show that AMU maintains cleaner and more retrievable personalized memories, and an SLM-guided (Small language model guided) structured framework for writing-time memory control.
This work learns a shared semantic frame and sparse coordinates that reconstruct semantic displacements while suppressing nuisance variation, with anchor-dependent diagonal modulation adjusting atom strengths without sample-specific rotations to support reusable invariant directions as a sparse coordinate system for lo...
A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for la...
Xuan-Yi Zhou, Qiu-Yang Mang, Huan-Zhi Mao et al.· 0 citations
Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation....
Wen-Bin Hu, Hui-Hao Jing, Hao-Chen Shi et al.· 0 citations
Background. Chest radiographs (CXRs) are the most common imaging exam performed, but CXR reports can be nonspecific due to incomplete history, relying on vague terms like opacities to convey diagnostic uncertainty. We prospectively evaluated whether introducing a clinical dashboard with relevant patient information aff...
J. Balkman, S. Chimmula, Alysha Lam et al.· medRxiv· 0 citations