Skip to content

Category

software testing

2,520 papers

AuraForge: Scaling Security Supervision for Training Coding Agents

Coding agents are now proficient enough to generate complex software applications from a single prompt. As their capabilities have grown, human oversight has increasingly shifted from line-by-line code review toward hands-off evaluation of outcomes. However, recent studies have shown that such a transition exposes a cr...

Dan-Qing Wang, Song-Wen Zhao, Harsh Sharma et al. · 0 citations

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, a...

Xuan Zhang, Long-Tao Zheng, Cun-Xiao Du et al. · 2 citations
#machine learning Preprint Oct 2026

TRACE: Tackling Real-World Resource Assignment Problems via Agentic Heuristic Design

Dynamic resource assignment, the real-time allocation of task streams to heterogeneous processing nodes, is the backbone of modern computing infrastructure. While learning-based schedulers excel in research, industrial deployments still rely on hand-written rules that operators can read, audit, and execute within tight...

J. Ayala-Romero, Andres Garcia-Saavedra, X. Costa-Pérez · 0 citations
#machine learning Preprint Oct 2026

FastCI: Efficient GPU-Intensive CI for LLM Training Frameworks

As large language models (LLMs) keep growing in size and complexity, their training frameworks evolve at a rapid pace as well. Therefore, continuous integration (CI) is critical for maintaining the quality and stability of these frameworks. However, unlike traditional software, CI for LLM training frameworks relies on...

Tian-Shuo Qiao, Nai-Qian Zheng, Xiao-Peng Liu et al. · 0 citations
#software testing Review Sep 2026

Innovation Shaping the Future: The Relationship Between Innovative Behavior and Vision About the Future Among Sports Management Students

The purpose of this research is to examine the relationship between innovative behaviors and future visions of sports management students. The research was conducted within the framework of a relational survey model and was carried out using data obtained from 215 students (90 female, 125 male) studying in the Eastern...

Gamze Durmuş, Pınar Karacan Doğan · 0 citations
#software testing Preprint Sep 2026

Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems

This work emulates a software-engineering workflow in which a planner represents a company hiring an external developer and reads the nondisclosure rule as banning plaintext, not character codes or riddles, and shows that benign agents can cross the same boundaries without adversarial incentives.

Deema Alnuhait, Geng-Yu Wang, Muhammad Khalifa et al. · 0 citations
#software testing Dataset Open access Sep 2026

SLEF Reference Evaluation Dataset v1.1

Reference evaluation dataset for SLEF (Synthetic Learner Evaluation Framework) — a controlled-Ground-Truth framework for evaluating learner-state inference under synthetic learner conditions. This release contains the scientific evidence produced by the SLEF Reference Evaluation across five tutor conditions: the synthe...

Marco Iannacone · 0 citations
#software testing Open access Sep 2026

P15-CROSSMAP: A fail-closed relation gate

Documentation-only v0.2.0 of P15-CROSSMAP. This original English technical note classifies three bounded relations among the locally tested R15 Maxwell fixtures: the conforming R5 oracle versus the hybrid beta=4 pencil, the historical beta=4 failure versus the beta=24 stabilization candidate, and source-problem refinem...

Riccardo Giudici · 0 citations
#software testing Open access Sep 2026

ramansep: two-mode separation of strain and carrier density in 2D-material Raman maps

Overlapping peaks, pixels without a peak, gaps in smoothed maps, and calibration uncertainty for any number of modes. Added fit_two_modes(..., joint=True) and fit_map(..., joint=True): the two peaks are fitted together (two Lorentzians on one shared constant or linear baseline) over the range spanned by both windows, s...

Tanvir M. Mahim · 0 citations
#software testing Open access Sep 2026

Reproducibility archive: Disparate Impact of Missing-Data Policy in Composite Healthcare Star Ratings

Reproducibility archive for the article Disparate Impact of Missing-Data Policy in Composite Healthcare Star Ratings: a pre-registered audit of three missing-data policies (the published CMS Overall Hospital Quality Star Rating design, median impute-then-rank, and an abstaining ranker whose uncertainty interval is buil...

Haitham A. El-Ghareeb · 0 citations
#software testing Dataset Open access Sep 2026

Dataset for "A Comparative Study of Different Deep Learning Techniques for Software Vulnerability Detection"

datasets Contains train, validation and test splits from 3 different datasets: * BigVul* MegaVul* ProjectKB results test_cases_and_predictions.csv Contains test case level data like CWE-ID, OWASP top 10 category, MVC category, fix size, fix type and also predictions from all three methods.Interpretation of ground_truth...

Anonymous, Lóránt Értekes, Rosmael Zidane Lekeufack Foulefack et al. · 0 citations
#software testing Open access Sep 2026

FCoP: A Filename-as-Protocol coordination layer for multi-agent AI development

FCoP (File-based Coordination Protocol) is a persistent, project-visible coordination protocol for multi-agent work. Core owns protocol facts, validation, and state transitions; the canonical MCP server is a thin adapter whose manifest defines the exposed tool surface. Runtime, Profile presets, Host configuration, and...

W M Zhu · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.