Post-training compression of LLM attention is often formulated as independent matrix approximation, ignoring both the shared structure among attention projections and the representation shift introduced by earlier compression. We propose FTC, a sequential structured compression framework that adapts the approximation t...
Jiang-Feng Chen, Xin-Yu Wang, Tian-Shuo Yan et al.· 0 citations
Compression reports summarize how far a compressed language model moved from the dense one, usually by a KL divergence; a deployment that relies on the dense model's outputs needs to know how many of its decisions changed. We show that total variation, not KL, answers this directly. Across 802 compressed and perturbed...
Knowing how much a causal predictor could improve need not reveal the gain of the repair actually learned. We quantify this gap in a scalar Gaussian causal experiment with known intervention geometry: auxiliary data identify effect magnitude up to bounded contamination, while diagnostics identify direction. The target...
Neural combinatorial optimization typically assumes a centralized solver that reads the whole instance. We study the opposite: combinatorial optimization under a hard information horizon, where every node commits to its share of a global solution seeing only its $k$-hop neighborhood, and those commitments must compose...
Johannes F. Loevenich, Thies Moehlenhof, Laurin Holz et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Open-ended query generation lacks standard answers, thus necessitating an effective reward mechanism. Pointwise scoring rubrics provide limited information about the relative quality of sample answers under the same prompt; merging multiple rubric judgments into a single score may also mask the differences between thes...
Zi-Hang Shen, Qi Liu, Zi-Xuan Yang et al.· 0 citations
Replay selectors often rank cached trajectories by format feedback, confidence, freshness, or response length, although cache-level correctness and downstream learner utility are distinct objectives. We formalize this selection-to-learning gap and introduce FAER as an auditable full-trajectory replay framework. Its tra...
Miaobo Hu, Shuhao Hu, Xiaobo Guo et al.· 0 citations
The partial area under the receiver operating characteristic curve (pAUC) is an important performance metric for binary classification that summarizes true positive rates within a specific range of false positive rates (FPRs). Classifiers that achieve high pAUC need to be obtained in many real-world applications such a...
Atsutoshi Kumagai, Tomoharu Iwata, Taishi Nishiyama et al.· 0 citations
Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, leaving the agent with no signal for ranking states. Temporal abstraction, which treats k e...
P. Dutenhefner, Dikshant Shehmar, Wagner Meira et al.· 0 citations
Replication materials accompanying “The marginal value of alternative data in credit screening: Evidence from Chinese digital lending” by Yuan Chen and Jiawei Xu. The package contains instructions for obtaining the source data, environment specifications, data processing and modelling scripts, parameter settings, rando...
Chen Yuan, Xu Jiawei· Zenodo (CERN European Organi...· 0 citations
Replication materials accompanying “The marginal value of alternative data in credit screening: Evidence from Chinese digital lending” by Yuan Chen and Jiawei Xu. The package contains instructions for obtaining the source data, environment specifications, data processing and modelling scripts, parameter settings, rando...
Chen Yuan, Xu Jiawei· Zenodo (CERN European Organi...· 0 citations
Pharmaceutical companies increasingly rely on Medical Science Liaisons (MSLs), which is a structured, science-led process of engagement rather than a sales role, to exchange clinical evidence with specialist healthcare professionals (HCPs). However, few empirical studies have explored the relationship between the perce...
Muhammad Usman, Muhammad Amin· Qlantic Journal of Social Sc...· 0 citations
Social determinants of health interventions can significantly improve health and reduce inequities in sub-Saharan Africa and the need for increased investment in social determinants approaches to achieve health equity is supported.
O. Sanni, A. E. Sanni· Journal of interventional ep...· 0 citations
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.