A self-referential retrosynthesis framework for explainable AI provenance forensics under a fixed-generator setting that leverages a jointly optimized encoder-decoder pair to implement a self-embedding mechanism that enables round-trip consistency verification.
Yi-Jie Lin, Ching-Chun Chang, Isao Echizen et al.· 0 citations
Motivated by the observation that safety signals from unsafe text processing can be transferred, safety-awareness representation transfer (SRT) is proposed, a lightweight direction-refinement method that mitigates cross-modal safety drift with a frozen MLLM backbone.
Tian-Qi Xiao, Shiyao Cui, Ming-Hao Zhang et al.· 0 citations
This work evaluates five LLMs as memory writers and two as executors across procurement, cybersecurity, and finance and introduces EAL-Bench, which measures how accurately persistent memory preserves evolving authorization state and whether errors propagate to downstream unauthorized actions.
Tommaso Cerruti, Mika Okamoto, Ansel Kaplan Erol· 4 citations· ⚡1
This work introduces Homomorphic Encryption-Aware Training (HEAT), a fine-tuning method that makes the per-nonlinearity iteration counts learnable, enabling them and the model weights to co-adapt during training.
Alessandro Zirilli, Davide Marincione, E. Kornaropoulos et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Safety properties assessed separately for Model Context Protocol (MCP) tool use and Agent2Agent (A2A) delegation need not describe behavior when one agent uses both. We measure one such behavior in a single controlled MCP-to-A2A configuration: a testbed drives a real-model host across a local MCP and a local A2A leg in...
To separate component effects from matcher rewards, CTIForge is built, whose deterministic validation layer can vary while extraction is held byte-identical, and which coincides with a roughly 2.8-fold increase in actions explicitly disputing entity type.
Through harness-policy co-evolution, SafeEvolve converts safety experience into an evolved runtime harness and improved policy behavior, and experiments show that SafeEvolve achieves a stronger safety-utility tradeoff than existing baselines.
ClaimReceipt, a claim-relative receipt specification and selective verifier that binds typed transaction evidence to a signed experiment manifest and returns PASS, INVALID, or INCONCLUSIVE per claim, is introduced.
This SoK examines 134 relevant papers on privacy concerns in motion patterns recorded by XR headsets, including how adversaries can obtain users'motion patterns, the inferences they can draw from them, and methods for protecting users and clarifies the state of XR motion privacy.
Azim Ibragimov, Alina Vasina, Uliana Polshcha et al.· 1 citation
GuidedBench, a novel benchmark comprising a curated harmful question dataset and GuidedEval, an evaluation system integrated with detailed case-by-case evaluation guidelines are introduced, ensuring reliable and reproducible evaluations.
Ruixuan Huang, Xun-Guang Wang, Zongjie Li et al.· 8 citations
This work identifies a phenomenon termed token-level memorization asymmetry through theoretical analysis of diffusion training dynamics and proposes Q-Skew, a quantile-weighted skewness-based indicator for membership inference on finetuned DLMs.
Sheng-Fang Zhai, Leo Marchyok, Yu-Ling Shi et al.· 1 citation· ⚡1
The Tri-Layer Sieve is presented, a middleware defense that sanitizes retrieved evidence through cross-embedding-space clustering with an independent judge model, structural filtering of trigger-payload artifacts, and LLM consistency verification, and exploits a key weakness of retrieval-stage poisoning.
Muhaimin Bin Munir, Akib Jawad Ononto, Nazia Shehnaz Joynab et al.· 0 citations