VR-FraudNet, a five-stage framework combining a time-conditioned spectral graph encoder, LightGBM triage and threshold routing, a schema-constrained rationale language model, a fixed deterministic verifier, and an isotonic probability mixer with split-conformal calibration, does not establish universal adversarial safe...
Md Sultanul Arefin Sourav, Evha Rozario, Md Ashiqul Islam et al.· Discover Artificial Intellig...· 0 citations
A controlled counterfactual audit of 80 recent English-language arXiv manuscripts shows that metadata perturbations can displace manuscripts and alter top-K shortlist membership, identifying a concrete reliability risk for LLM-assisted scientific evaluation pipelines.
Marco Rospocher· Discover Artificial Intellig...· 0 citations
The results show that harness-managed control flow can substantially improve the effectiveness of the smallest models and suggest that small models can be useful for repository repair when responsibilities are divided across explicit stages that can be independently assigned to the component best suited to each.
Francesco Dente, Dario Satriani, Donatello Santoro et al.· 0 citations
Deploying Large Language Models (LLMs) on memory-constrained edge servers to serve requests from mobile devices is challenging due to their substantial resource demands. The Key-Value (KV) cache and Feed-Forward Network (FFN) parameters consume the majority of available memory. However, existing methods typically rely...
Zhong-Xiang Wei, Yi-Peng Zhou, Jin Zhao et al.· IEEE Transactions on Mobile...· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
OrthoPurify is proposed, a more efficient method to purify backdoored model weights via one-step orthogonal projection, which reduces the attack success rate to near zero while preserving the original performance across diverse benchmarks, without retraining the backdoored model or introducing inference-time overhead.
Bo-Jun Yang, Hao-Chen Zhou, Zhi-Fang Zhang et al.· 0 citations
This work constructs a signal by force-decoding a fixed set of candidate answers and develops a framework that manipulates the image and question text separately, making the boundary searchable for any behaviour formulated as a choice between candidate answers.
Linking people's appearance and actions to character identities is essential for understanding video narratives. We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation. St...
A. F. Razzouki, Killian Steunou, Khalil Guetari et al.· 0 citations
This work contributes a phonetically- and gender-balanced two-tier Bangladeshi Bangla corpus balanced via a tiered Jensen-Shannon divergence objective over conjunct clusters (juktakkhor), together with three fine-tuning changes: a merge-consistent tokenizer extension, Bangla text normalization, and a prompt-masked dual...
Emtiaz Uddin Ahmed, Araf Mahmud, S. Hossain et al.· 0 citations
A theory built on two quantities of the image interface: S, the number of visual tokens across an object's side, and L, the content a call must cover, which concludes that recognition improves gradually with S, and any search strategy, zoom agents included, obeys a recall-cost frontier.
Many annotation projects begin before experts have a stable guideline or enough labels to train a task-specific model. We present Goldsmith, an agentic pipeline that turns a small gold set---expert-annotated calibration examples representing the intended task boundaries---into a reusable structured annotation definitio...
Yi-Han Li, Han-Yi Zhang, Xiao-Xi Jiang et al.· 0 citations
Embedding tables are among the largest components of modern language models. Most compression methods fix a coding geometry such as coordinate blocks, low-rank subspaces, or unrestricted codebooks, and optimize within it. We instead ask whether the coding geometry can itself be discovered. We introduce \emph{OrBIT}, a...