Deep learning (DL) techniques are increasingly integrated into traditional software systems, giving rise to hybrid AI-enabled systems that combine neural models with program logic. While these systems exhibit remarkable capabilities, their complex and heterogeneous architectures pose significant challenges for reliabil...
Xin-Yu Gao, Yang Feng, Yu-Chen Lu et al.· 0 citations
Test driven development (TDD) is a controversial and interesting approach to software development; while many think of better tests as a primary purpose of TDD, in practice the goal is as much to use tests to encourage continued progress in coding. That goal however rests on the notion that TDD ensures tests are good e...
A human-LLM workflow that pairs each curated semantic mutant with an instructor-approved seed phrase for an on-demand LLM expansion, which shows that the curated semantic-mutant set contains a fraction of the mutants a traditional mutation engine produces.
Rebecca Williams Earle, Jonathan Bell· Proceedings of the 2026 ACM...· 0 citations
SAT solving is a core computational challenge across hardware verification, electronic design, and planning; an FPGA can evaluate SAT candidates at hardware clock rates without software overhead, enabling sub-microsecond solving on small instances. Yet, the throughput-area tradeoffs of parallel FPGA SAT datapaths remai...
Andrew Bonilla· Companion Proceedings of the...· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This proposed dissertation investigates mutation testing from three complementary perspectives: how mutations interact with program executions and test oracles to produce mutant kills and reveal opportunities for test improvement, and how mutation testing is adopted, configured, sustained, and acted upon in popular ope...
Hang Du· Companion Proceedings of the...· 0 citations
This talk explores how agent-first development platforms like Google Antigravity shift the paradigm from manual test authoring to autonomous software verification, and describes recent advances in autonomous visual testing.
Paige Bailey· Companion Proceedings of the...· 0 citations
PyMut4SE is a novel mutation tool for Python that focuses on a comprehensive set of mutations for any Python project, providing a rich and extensible set of mutation operators, access to mutated source code and its characteristics, detailed execution and behavioral observations, and support for both selective mutation...
Laura Plein, Matthieu Jimenez, Mike Papadakis· Companion Proceedings of the...· 0 citations
Memory-safety bugs are one of the oldest and most common sources of security vulnerabilities, and their modern-day prevalence is a consequence of the widespread use of non-memory-safe languages such as C and C++. Transitioning away from C and C++ to memory-safe languages, namely Rust, to eliminate memory-safety bugs is...
Victor Chen· Companion Proceedings of the...· 0 citations
Fuzzing finds bugs by testing software behavior with high volumes of input. To create semantically valid inputs for different targets, traditional fuzzers require significant engineering effort. LLM-based fuzzing approaches can be adapted to create inputs for a wide range of software but introduce new issues: every tes...
Lu Maltsis· Companion Proceedings of the...· 0 citations
The paper asks whether high-level conceptual specifications improve LLM-generated PBT quality and whether they help developers extend and maintain AI-generated systems (RQ2), and reports preliminary results applying Spinach to two open-source applications.
Savitha Ravi, Michael Coblenz· Proceedings of the 1st Inter...· 0 citations
LLM-based testing can expose subtle reliability issues in numerical software, but monolithic prompts make it hard to inspect which guidance drives effectiveness. We study whether Agent Skills can make such workflows more explainable by factoring procedural knowledge into composable testing components. Using numerical i...
Yu-Tong Wang, Cindy Rubio-González· Proceedings of the 2nd Works...· 0 citations
Specifying contracts as program behavior using Hoare triples is an established practice in software engineering. Compared to traditional testing approaches, formal verification of these contracts provides stronger guarantees of functional correctness. However, its practical adoption has remained limited since significa...
Aryan Kumar, Alex Toppo, Sandip Ghosal· Proceedings of the 2nd Works...· 0 citations
Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.
Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…
AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.