Skip to content

Category

artificial intelligence

14,156 papers

#artificial intelligence Preprint Open access Oct 2026

Agent4RE: A Self-Refining Multi-agent Framework for End-to-End Software Requirements Engineering and Benchmarking

Existing LLM-based approaches for software Requirements Engineering (RE) typically rely on basic prompting strategies or rudimentary agent collaboration, under-utilizing the full potential of multi-agent systems. Meanwhile, available datasets focus on isolated subtasks, such as requirements extraction, classification,...

Yongjian Tang, Linhan Li, Thomas Runkler · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SAVU-BENCH: A Real-World Benchmark for Spatial Audio-Visual Understanding

Spatial audio-visual understanding requires models to recognize not only what is present, but also where events occur and how they relate across modalities. Existing benchmarks often rely on simulated scenes, evaluate isolated spatial skills, and provide limited diagnostic insight into failure modes. We introduce SAVU-...

Yu Chen, Ruihang Liu, Yangguang Xu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

WorldBench: Evaluating LLMs on Three.js Voxel World Generation

Large language models can now write complete, interactive 3D worlds as code, but grading those worlds automatically is unreliable. Existing judges take one view of the output: a vision-language model scores a few rendered snapshots, or a language model reads the source. On worlds written by five frontier models we find...

Krish Bakshi · 0 citations
#artificial intelligence Preprint Open access Oct 2026

When AI Finds Hidden Messages, Does It Report?

When an assistant encounters a message for another AI, does it tell its user? Four fixed model-provider deployments perform simulated source tasks in 1,280 ordinary-note and 128 enhanced-note sessions. Harmless and harmful messages have matched plaintext and ROT13 versions, with no-message controls. Observers receive n...

William Guey, Rashik Jahangir, Pierrick Bougault et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution

Large language model (LLM) agents are rapidly reshaping software engineering, accompanied by an explosion of new code benchmarks. Yet nearly all existing benchmarks still rely on the same decades-old criterion: a solution is correct if it passes a fixed set of unit tests. Such tests are often insufficient: they check o...

Shuangjie Yao, Hao Wang, Koushik Sen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking

In post-deployment time, inputs to deep learning models may or may not be adversarially patched. Patch robustness certification on such inputs within a patch bound can verify their label benignity and should retain high prediction accuracy. However, existing smoothing-based and masking-based recovery defenders cannot a...

Qilin Zhou, Zhengyuan Wei, Haipeng Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Code Understanding is a Bottleneck for Coding Agents

Repository benchmarks (e.g., SWE-bench) for coding agents often assume that lines of code edited can predict task difficulty, but such datasets' poor control over code and task types makes it hard to know which abilities truly drive agent errors. We present CABRA: a Coding Ability Blueprint for Rigorous Agent evaluatio...

Nishant Balepur, Kiran Tomlinson, Tobias Schnabel · 0 citations
#artificial intelligence Preprint Open access Oct 2026

From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage

Security operations centers (SOCs) must triage large volumes of alerts, most of which are benign, while missed attacks can remain uninvestigated. Tool-using large language model (LLM) agents can retrieve evidence during triage, but it remains unclear how reasoning strategies determine what to gather and when an investi...

Saimon Amanuel Tsegai (Daphne), Alex Kantchelian (Daphne), Danfeng (Daphne) et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Beyond Type-checking: Towards Holistic Evaluation of Formal Specification Generation

When generating verifiable code, natural language requirements are mapped to machine checked code using LLMs and agentic workflows. A crucial component of this pipeline is specification generation (SpecGen), which produces a formal contract against which an agent can prove implementation correctness. Proof generation c...

Srijith Nair, Aditya Vempaty, Jia Liu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging

Public leaderboards for AI models are read continuously, and attackers can see every published standing. Vote rigging, selective disclosure of private variants, and benchmark contamination can each move a ranking. Existing guarantees assume genuine records or bound the corruption per step, which an attacker who corrupt...

Hamed Khosravi, Xiaoming Huo · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Survey on LLM-Integrated Hardware Design Verification

Large language models (LLMs) are increasingly being integrated into hardware verification to automate specification interpretation, verification-artifact generation, debugging, formal reasoning, and tool orchestration. This survey provides a systematic review of LLM-assisted hardware functional verification across Syst...

Hao Zheng, Jaime Rafael Imperial, Bardia Nadimi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows

Multimodal large language models (MLLMs) often answer visual reasoning questions by relying on linguistic priors rather than task-relevant visual evidence. Textual chain-of-thought reasoning can partially mitigate this issue by encouraging models to decompose visual questions into intermediate evidence-seeking steps, b...

Albert Gao, Bing Xue, Andrea Zanette · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.