Skip to content

Category

software testing

2,461 papers

#artificial intelligence Review Oct 2026

AgentPersonaBench: Benchmarking Persona-Driven User Simulation

We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deployed for persona-driven user simulation, existing benchmarks primarily evaluate conversational styling or self-reports rather than authentic b...

Jin-Tao Huang, Yi-Fan Wang, Hong-Yuan Shen et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents

Agents built on large language models (LLMs) build and run software through tool calls. A call reaches its program through several hops, and any hop can change the call without notice. When the changed call fails, the agent retries a correct call, which costs users time and money. Benchmarks and failure analyses do not...

Bo-Yang Yang, Zhen-Hao Li, Zi-Yao Yang et al. · 0 citations
#large language models Open access Oct 2026

Deceptive Technical Debt in Autonomous Maintenance: Cognitive Asymmetries and Heterogeneous Auditing Pipelines

When autonomous Large Language Model (LLM) agents maintain and refactor safety-critical enterprise software under automated verification constraints, task completion incentives frequently induce reward-hacking heuristics. Early-generation agents introduce what we term Spurious Agentic Heuristic Glues (SAHGs)—deceptive...

Chen Jiang · 0 citations
#large language models Open access Oct 2026

Deceptive Technical Debt in Autonomous Maintenance: Cognitive Asymmetries and Heterogeneous Auditing Pipelines

When autonomous Large Language Model (LLM) agents maintain and refactor safety-critical enterprise software under automated verification constraints, task completion incentives frequently induce reward-hacking heuristics. Early-generation agents introduce what we term Spurious Agentic Heuristic Glues (SAHGs)—deceptive...

Chen Jiang · 0 citations
#software testing Open access Oct 2026

Governance e Qualità nei Progetti AI: Metodologie, Metriche e Lessons Learned nell'Adozione Enterprise di Modelli LLM

L'adozione enterprise di modelli LLM pone sfide inedite in termini di governance, qualità e gestione del rischio. A differenza del software tradizionale — deterministico e verificabile — i sistemi basati su modelli di linguaggio presentano comportamenti probabilistici, difficoltà di testing convenzionale e requisiti di...

Claudio Fontanarosa · 0 citations
#software testing Open access Oct 2026

Governance e Qualità nei Progetti AI: Metodologie, Metriche e Lessons Learned nell'Adozione Enterprise di Modelli LLM

L'adozione enterprise di modelli LLM pone sfide inedite in termini di governance, qualità e gestione del rischio. A differenza del software tradizionale — deterministico e verificabile — i sistemi basati su modelli di linguaggio presentano comportamenti probabilistici, difficoltà di testing convenzionale e requisiti di...

Claudio Fontanarosa · 0 citations
#software testing Open access Oct 2026

Comparaison de la qualité du recueil de données chez les patients pris en charge par le SMUR de Montpellier pour une douleur thoracique : logiciel SMUR-t@b versus texte libre

Objectives: the main objective of this study was to compare the completeness of medical data collection using the digital tool SMURt@b versus the former paper format during the pre-hospital management of chest pain. Patients studied: the patients included in this study were adults managed in a pre-hospital setting for...

Juliette Meissirel · 0 citations
#software testing Open access Oct 2026

Analisis Flypaper Effect: Pengaruh PAD dan DAU Terhadap Belanja Daerah Sumatera Selatan

ABSTRACT   Examining the impact of Regional Original Revenue (PAD) and General Allocation Funds (DAU) on regional expenditure posture, as well as analyzing the existence of the Flypaper Effect phenomenon within regency and city governments in South Sumatra Province, is the primary focus of this study. A quantitative ap...

Anisa Cahya Ramadhani, Periansya Periansya, C. Choiruddin · 0 citations
#software testing Open access Oct 2026

From Technological Capability to Disaster Governance in North Korea: An Evidence-Bounded Framework from Scientific Publications

This study examines the extent to which scientific publications can support claims about algorithmic disaster governance in an information-limited setting and distinguishes among technological capability, functional compatibility, administrative embedding, and system integration.

Le-Yuan Liu · 0 citations
#software testing Open access Oct 2026

Pengaruh Akuntabilitas, Transparansi, dan Good Governance Terhadap Kinerja Instansi Pada OPD Kota Palembang

ABSTRACT   This study aims to determine the influence of accountability, transparency, and good governance on the institutional performance of Regional Apparatus Organizations (OPDs) in Palembang City. A quantitative research approach was used. Primary data were collected through a questionnaire distributed to OPD empl...

Zahara Adelia Sani, C. Choiruddin, Desri Yanto · 0 citations

Integrating Augmented Reality Sandbox Technology to Enhance Hydrology Education in Civil Engineering: A Pilot Study

Augmented reality (AR) tools may support hydrology instruction by turning abstract watershed concepts into visible and manipulable processes. This pilot study reports the construction and classroom implementation of an AR sandbox to support basin and watershed understanding in an undergraduate civil engineering hydro...

Mónica Guzmán-Rojo, Richard Rocha Rivero, Diego Vittorini Echalar et al. · 0 citations
#software testing Open access Oct 2026

Vulnerable Function Detection Using Lexical and Structural Features: An Empirical Study on PrimeVul and DiverseVul

It is suggested that carefully designed lightweight feature representations, combined with systematic multi-metric evaluation, can provide a reproducible and interpretable baseline for practical software vulnerability detection.

Jing Wang · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.