Skip to content

Output Format x Model Identity: Interaction Effects in Single-Round Coding Agent Performance

Jul 2026 · arXiv.org · Vol abs/2607.21674 · 0 citations · 15 references
Computer Science

TL;DR

A model-specific output strategy, a tool-design principle that constrains format semantics to the agent's own localization step, and release all data and templates for reproducibility are proposed.

Abstract

Output format is not a neutral implementation detail -- it can reorder model rankings, amplify or suppress individual model differences, and determine whether a coding agent succeeds or fails. We conducted a controlled single-round experiment with 3 models (DeepSeek V4, Doubao 2.0 Pro, Qwen 3.7 Max) x 3 output formats (full file, JSON Patch, unified diff) x 6 tasks x 20 repetitions, totaling 4,013 runs across 4 open-source projects. Only one project (tqdm) yielded non-zero success rates: dotenv, requests, and jsoup yielded zero successes in 2,551 runs. Our central finding is a format x model interaction with no universally optimal format. Doubao achieves 94% success with JSON Patch (Cohen's h = 1.57, p<0.001), DeepSeek excels at unified diff (66%, h = 0.63), and Qwen shows a small but significant full-file preference (50%, h = 0.29, p<0.05). Beyond these headline results, we identify a distinct failure mechanism -- format misuse -- where agents correctly diagnose a problem but execute it with excessive scope, most vividly when a one-line fix is applied as a full-file replacement. We propose a model-specific output strategy, a tool-design principle that constrains format semantics to the agent's own localization step, and release all data and templates for reproducibility. All experimental data are available at https://doi.org/10.5281/zenodo.21505157.

View source

Similar papers

#natural language process... Preprint Sep 2026

Playing log(N)-Questions over Wikipedia Abstracts: How Per-Round Errors Compound Under Information Asymmetry

We evaluate six frontier language models on the two-agent $\log_2 N$-Questions game (Potash et al., 2019) to measure self-communication across an information asymmetry. A questioner with access to $N$ candidate Wikipedia lead paragraphs ($N = 4$ to $1024$) must identify a secret target using exactly $\log_2 N$ binary q...

P. Potash · 0 citations
Preprint Aug 2026

Architecture as Capability Equalizer for Coding Agents

LLM-based coding agents generate complete software systems from high-level descriptions, yet little is known about how the format of architecture specifications affects the quality of generated code or whether this effect depends on model capability. We present a controlled experiment comparing five informationally equ...

A. Canedo · 1 citation
Preprint Sep 2026

ToMAS: A Pilot Failure-Grounded Theory-of-Mind Benchmark from Multi-Agent LLM Failures

LLM-based multi-agent systems can fail even when communication succeeds because agents do not correctly track their peers'roles, knowledge, or intentions. We investigate whether such inter-agent misalignment cases, labelled FC2 in MAST-Data, can be converted into functional partner-state reasoning items. ToMAS applies...

M. Ishfaq, Glaucia Melo · 0 citations
Jul 2026

Moral Hazard in Multi-Agent Language Models

CREDIT (Counterfactual Replay for Evidence-Driven Information Transfer), a mechanism-aligned multi-agent prompt-optimization algorithm that uses matched hidden-state twins and total-action replay to reward robust causal contribution rather than query frequency is introduced.

Dane Malenfant · 0 citations
#artificial intelligence Preprint Sep 2026

Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead

Small differences on coding-agent leaderboards are often read as an ordering of systems. We audit whether the published verdicts support this reading, using 254 SWE-bench submissions across four splits without running models. On Verified, the leading two entries each resolve 396 of 500 instances. The top ten share 285...

Feng-Shuo Liu, Ying Liu, Rui-Ze Sun et al. · 0 citations
#artificial intelligence Conference Open access Sep 2026

Self-Reports Are Not Verification

An environment-grounded audit is introduced in which every intermediate proposal receives an exact outcome in an evolutionary Contexto search whose feedback function assigns every valid guess an exact rank without human annotation.

En-Rong Pan, Ryan Zhou, Ting Hu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.