Skip to content

Category

software testing

2,461 papers

#software testing Preprint Oct 2026

Autoware in Construction: Gap Analysis and LiDAR Perception Toward Off-Road Autonomous Driving

Autoware is an open-source autonomous driving software platform widely adopted by researchers and industry developers. Originally developed primarily for public-road applications, including passenger vehicles, taxis, and buses, Autoware is increasingly being extended to off-road environments such as construction and ag...

Yu Otsuki, Teja Emmey, Sena Matsushita et al. · 0 citations
#machine learning Preprint Oct 2026

Correct Verdicts, Flawed Reasoning: Structured Auditing of LLM-based Vulnerability Reasoning

Large Language Models (LLMs) are increasingly deployed for automated software vulnerability analysis. Binary classification alone is insufficient; practitioners need explanations to triage bugs and engineer patches. Standard practice relies on Chain-of-Thought (CoT) prompting, but free-form reasoning allows models to o...

Boyue Caroline Hu, K. Ahir, Ronghao Ni et al. · 0 citations
#machine learning Preprint Oct 2026

Mind the Gaps: From Failure Attribution to Closed-Form Repair of Code Language Models

Code language models must be maintained like the software around them: when a library evolves, a model keeps writing the interface that it saw during training. Repairing the model itself lets one correction reach all downstream uses. Existing repair methods attribute a failure to neurons, select the highest-ranked ones...

Jian Gu, Hong-Yu Zhang, Chun-Yang Chen et al. · 0 citations
#artificial intelligence Review Oct 2026

Does AI Help Cyber Attackers or Defenders? Evidence from Nonpublic Vulnerabilities and Subsequent Attacks

The release decision for frontier AI systems increasingly relies on cyber capability benchmarks, yet public vulnerability benchmarks can expose agents to previously published advisories, exploits, and fixes, making it difficult to distinguish prior exposure from capability on unseen vulnerabilities. We evaluate open-we...

Tobias Heldt, Matthew Turk, Christoph R. Landolt et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Correct Code, Broken Contributions? SWE-CC: Benchmarking Repository Policy Compliance for Coding Agents

Autonomous coding agents now resolve a substantial share of real-world GitHub issues. However, passing functional tests differs fundamentally from producing a high-quality contribution acceptable for merging. Mature open-source projects publish repository-specific contribution policies, spanning style, git, testing wor...

Hải Đăng Trương, Rayner Goh, Thanh Le-Cong et al. · 0 citations
#artificial intelligence Preprint Oct 2026

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery. We introduce UndoBench, a benchmark spanning 36 base workflows and 36 fault scenarios across...

Dolly Sah, Tanmay Sah, Harshul Jain et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Software World Models: From Consequence Prediction to Decision Value

A coding agent may safely modify one repository while silently breaking downstream services, libraries, or datastores that depend on it. Exhaustively running integration tests after every agent action is impractical, so the agent must predict these failures before executing them. Existing software world models predict...

Tong-Li Su, Yun-Tong Hu, Liang Zhao et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Complex Agents, Shallow Tests: Demystifying and Enhancing Test Adequacy of Agent Harness in the Wild

LLM-based agentic systems are emerging as a new software paradigm. Modern agents are typically composed of backbone LLMs and a surrounding harness that serves as the operational software infrastructure for agent execution. As agent harnesses grow increasingly complex, agents suffer from diverse harness implementation b...

Yi-Fan Xiong, Jing-Yi Ge, Zhen-Peng Chen et al. · 0 citations
#artificial intelligence Preprint Oct 2026

TeleTune: Evolving Agent Skills From Offline Telemetry

Computer-use agents need to capture procedural knowledge of how people use software. User telemetry offers a scalable source of this knowledge. However, learning reusable skills from these logs requires addressing three challenges: (1) Goal Underspecification, since logs do not record the goal behind each action; (2) N...

J. Chen, Elias Stengel-Eskin, Yan Chen et al. · 0 citations
#artificial intelligence Preprint Oct 2026

How corner is a corner case? Percentile control for highway scenario generation

Generating corner-case scenarios with appropriate adversity in a simulation environment is critical for testing an autonomous vehicle (AV) software stack's safety performance before deployment. Existing autonomous-driving scenario generators can enforce specific behavior, adversity, or feasibility conditions, but they...

Jia-Xi Liu, Hang Zhou, Hang-Yu Li et al. · 0 citations
#artificial intelligence Preprint Oct 2026

MemTrace: State-Consistent Memory for Long-Horizon Coding Agents

As coding agents take on long-horizon software evolution tasks spanning multiple files and stages, longer execution trajectories introduce two coupled challenges: (1) accumulated histories strain context budgets, and (2) repository changes can invalidate earlier execution evidence. Existing approaches address these cha...

Hong-Ming Xu, Le Zhou, Zhong-He Jin et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Recursive Improvement of a Differentiable Scientific Software Ecosystem

Differentiable programming connects scientific computation with gradient-based inference, learning and design. Extending these capabilities across a heterogeneous software ecosystem requires specialized effort to implement derivatives, integrate interfaces and evaluate quality. AI coding agents can accelerate this tran...

Peng-Cheng Hou, Xiao-Jun Tan, Si-Han Hu et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.