Current multimodal software-engineering benchmarks expose images as additional context, but do not test whether an agent can integrate evidence distributed across images into a verified repository-level repair. We present SWE-PolyVision, an executable benchmark of 92 real tasks from 36 open-source organizations, with 4...
Jia-Jun Wu, Lei-Xin Sun, Zi-Hang Tan et al.· 0 citations
SWE-Prometheus is presented, a benchmark for the broader task of improving repository engineering governance, showing why repository-governance evaluation should report improvement, behavior preservation, evidence quality, and coverage together.
Jia-Jun Wu, Lei-Xin Sun, Zi-Hang Tan et al.· 0 citations
This work argues RSI must next operate across real, diverse scientific, engineering, and meta-scientific domains, not where formal evaluation is merely tractable, and presents MetaRSI-v1, where improvement is the scheduled composition of three typed operators over one unified paradigm.
Zi-Hang Tan, Lei-Xin Sun, Zi-Tong Shi et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.