We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic b...
Jin-Tao Huang, Yi-Fan Wang, Hong-Yuan Shen et al.· 0 citations
Agents built on large language models (LLMs) build and run software through tool calls. A call reaches its program through several hops, and any hop can change the call without notice. When the changed call fails, the agent retries a correct call, which costs users time and money. Benchmarks and failure analyses do not...
Bo-Yang Yang, Zhen-Hao Li, Zi-Yao Yang et al.· 0 citations
When autonomous Large Language Model (LLM) agents maintain and refactor safety-critical enterprise software under automated verification constraints, task completion incentives frequently induce reward-hacking heuristics. Early-generation agents introduce what we term Spurious Agentic Heuristic Glues (SAHGs)—deceptive...
Chen Jiang· Zenodo (CERN European Organi...· 0 citations
When autonomous Large Language Model (LLM) agents maintain and refactor safety-critical enterprise software under automated verification constraints, task completion incentives frequently induce reward-hacking heuristics. Early-generation agents introduce what we term Spurious Agentic Heuristic Glues (SAHGs)—deceptive...
Chen Jiang· Zenodo (CERN European Organi...· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
L'adozione enterprise di modelli LLM pone sfide inedite in termini di governance, qualità e gestione del rischio. A differenza del software tradizionale — deterministico e verificabile — i sistemi basati su modelli di linguaggio presentano comportamenti probabilistici, difficoltà di testing convenzionale e requisiti di...
Claudio Fontanarosa· Zenodo (CERN European Organi...· 0 citations
L'adozione enterprise di modelli LLM pone sfide inedite in termini di governance, qualità e gestione del rischio. A differenza del software tradizionale — deterministico e verificabile — i sistemi basati su modelli di linguaggio presentano comportamenti probabilistici, difficoltà di testing convenzionale e requisiti di...
Claudio Fontanarosa· Zenodo (CERN European Organi...· 0 citations
Objectives: the main objective of this study was to compare the completeness of medical data collection using the digital tool SMURt@b versus the former paper format during the pre-hospital management of chest pain. Patients studied: the patients included in this study were adults managed in a pre-hospital setting for...
Juliette Meissirel· INRIA a CCSD electronic arch...· 0 citations
ABSTRACT
Examining the impact of Regional Original Revenue (PAD) and General Allocation Funds (DAU) on regional expenditure posture, as well as analyzing the existence of the Flypaper Effect phenomenon within regency and city governments in South Sumatra Province, is the primary focus of this study. A quantitative ap...
Anisa Cahya Ramadhani, Periansya Periansya, C. Choiruddin· Jurnal Media Akuntansi (Medi...· 0 citations
This study examines the extent to which scientific publications can support claims about algorithmic disaster governance in an information-limited setting and distinguishes among technological capability, functional compatibility, administrative embedding, and system integration.
ABSTRACT
This study aims to determine the influence of accountability, transparency, and good governance on the institutional performance of Regional Apparatus Organizations (OPDs) in Palembang City. A quantitative research approach was used. Primary data were collected through a questionnaire distributed to OPD empl...
Zahara Adelia Sani, C. Choiruddin, Desri Yanto· Jurnal Media Akuntansi (Medi...· 0 citations
Augmented reality (AR) tools may support hydrology instruction by turning abstract watershed concepts into visible and manipulable processes. This pilot study reports the construction and classroom implementation of an AR sandbox to support basin and watershed understanding in an undergraduate civil engineering hydro...
Mónica Guzmán-Rojo, Richard Rocha Rivero, Diego Vittorini Echalar et al.· Journal of Civil Engineering...· 0 citations
It is suggested that carefully designed lightweight feature representations, combined with systematic multi-metric evaluation, can provide a reproducible and interpretable baseline for practical software vulnerability detection.
Jing Wang· Journal of Cyber Security an...· 0 citations
Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.
Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…
AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.