Listen and ask questions on Gemini Notebook: https://notebook.google.com/notebook/31300077-8066-4826-bca1-faf9fe8dbd2e?authuser=1 The Little Twist in the Middle: History, Accessibility, Logos, and the Problem of Making the Other Reachable develops a unified account of how accumulated history changes what becomes access...
Ryan MacLean· Zenodo (CERN European Organi...· 0 citations
Background: Young people spend hours on social media every day, where following influencers is the norm. Influencers post frequently about alcohol and alcohol brands and there has been a rise in influencer marketing. However, prevalence estimates rely on manual analyses of only a small number of posts. We used a large...
Jack Delmenico, Dan Anderson‐Luxford, Emmanuel N. Kuntsche et al.· 0 citations
Comparing three model families and multiple math benchmarks, COMPASS outperforms the activation-steering baselines the authors compare against, improves GSM8K accuracy by 16 percentage points on average, and approaches CoT accuracy with 20-70\% fewer generated tokens.
Pratyay Dutta, Kowshik Thopalli, V. Narayanaswamy· 0 citations
Advisory systems built on large language models should be audited at the disclosure levels users actually reach, and judged across the whole identity space rather than one attribute at a time.
This work introduces BOTTLED, a benchmark in which agents receive an entire unlabelled workload and must complete it under fixed time, compute and LLM API budgets, and finds that strong zero-shot task performance does not reliably translate into strong bottling capabilities.
Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein et al.· 0 citations
This study evaluates the performance of lightweight open-source language models to resolve an input data entity against its corresponding best fitting SDM representation under resource-constrained conditions and provides significant and valuable insights into model selection, task formulation, and deployment strategies...
Cristian Martella, A. Martella, Antonella Longo et al.· 0 citations
System-level results provide system-level evidence that deployment-time safety depends not only on the capability of individual guardrails, but also on how their decisions are organized and coordinated.
Xing-Ru Zhou, Luis Sentis, Aarti Choudhary· 0 citations
DynaCore is presented, a unified architecture for efficient LLM serving via system-architecture co-design that substantially reduces service-level latency over quantization and reconfigurable accelerators, and proposes disaggregated quantization, applying dual-side quantization to prefill and weight-only quantization t...
Cong Guo, Chi-Yue Wei, Bo-Wen Duan et al.· 0 citations
This work presents HarnessSecurity, the first systematic empirical study and benchmark of open- and closed-source coding agent harnesses, and derives a ten-mechanism taxonomy and assesses 400 harness-mechanism cells using independent ratings by researchers and large language model judges.
Zheng-Yang Zhu, Li-Ming Huang, Run-Min Ji et al.· 0 citations
This paper finds that the emergence of massive activations is controlled by a single channel in the input embedding to a spike feed-forward network (FFN), and names this channel the massive activation gating channel (MAGC).
Min-Jia Mao, Shi Chen, Bo-Wen Yin et al.· 0 citations
This case study documents a CRA preparedness pilot for one such product, SEUXDR, an AI-augmented security monitoring product with a large-language-model active-response component on the open-source CYBERFORT platform, and offers practitioners a replicable starting point for translating CRA legal text into operational p...
Georgios Koutidis, Nikolaos Kekatos, Marina Korgiala-Karyda et al.· 0 citations
SCSM constructs pairs of pretraining views from the same group of unlabeled traces through Segmentation, Combination, Scaling, and Masking, which produces diverse observable patterns while preserving the underlying packet events and local traffic dynamics of real trace fragments.
Xian-Wen Deng, Rui-Jie Zhao, Ming-Wei Zhan et al.· 0 citations