It is concluded that instability should be treated as a standard evaluation axis in SE optimization, which should be routinely measured, reported alongside performance, and used to calibrate trust in any single run.
Abstract
In software analytics, rerunning the same analysis twice often yields different models and conclusions. This reduces trust in the model and limits its use. We find that model instability is a major problem. Across 127 multi-objective SE optimization problems (12,700 test cases), repeated runs of a state-of-the-art optimizer agree on only 13.7% of test cases, even under improved settings. We argue that this instability is not merely noise to tolerate, but a property that can be measured and managed. By adjusting how labels are spent, how complex the models become, and how splits are scored, we obtain models that agree 4.8 times as often as the default configuration. The standard deviation of optimization error falls by 22% on average (mean std 17.4 to 13.6), while recommendation quality improves rather than degrades. In terms of quality, the refined settings are statistically top-ranked on 119 of 127 datasets, compared to 74 for the defaults. We then test causal and data-locality interventions and find that they help only partially, suggesting a residual stability floor. Our evidence suggests there are fundamental limits to stability set by the data itself (noise, scarce labels, proxy objectives, and the many near-equivalent models a dataset admits). We conclude that instability should be treated as a standard evaluation axis in SE optimization, which should be routinely measured, reported alongside performance, and used to calibrate trust in any single run. The methods in this paper provide a baseline against which future efforts to reduce SBSE instability can be judged. To support open science, we offer the following reproduction package: https://tinyurl.com/Model-Instability
A graded, multi-family, failure-aware framework for stress-testing reasoning models that exposes structure that an aggregate score hides and exposes consistent weaknesses across all models.
Trusted monitoring has a cheap, trusted model score a stronger untrusted model's actions, and a diverse ensemble of them beats a single stronger monitor at matched cost. They are built by minimising average pairwise correlation, and that paper's twelve monitors shared one base model, leaving open what supplies the diversity. We study 24 open-weight monitors spanning nine pretraining lineages and a 29x range of detection skill (pAUC at 10 percent FPR, 0.028 to 0.803) on backdoored code. The metric used to build panels does not predict what a panel is for, and we can say why. Agreement on attack items splits into a shared-detectability signal component and an idiosyncratic error component, which predict ensemble gain with opposite sign (Spearman -0.25 and +0.26), so their sum, the metric actually used, predicts it barely at all (+0.05); the cancellation holds in 7 of 8 evaluations. Skill acts on signal (+0.53) while error stays flat (-0.01), which is why a monitor's own skill predicts its agreement with the pool (Spearman 0.84, n = 24, permutation p below 0.0001). Pretraining lineage is the obvious way to buy decorrelation, and it does not pay. At matched member capability, cross-lineage panels detect no better (permutation p = 0.13), and lineage barely moves the metric either (+0.064, p = 0.18). We report that against ourselves: on our own 22-monitor pool the same test read +0.104 at p = 0.037 until two monitors were added. An earlier pool topping out at pAUC 0.23 had already invalidated another analysis. Such a quantity is a property of the pool assembled. Panel gain over the best member falls monotonically with panel skill (-0.66 at k = 2, -0.70 at k = 3), and no correlation-weighted selection beats picking the single best monitor out of sample. Across six attacker models the gain result holds in all six, the agreement and cancellation results in five of six.
Standard evaluation of large language models is challenged by varying the token generation budget, i.e., the maximum tokens a model may produce, across seven levels, evaluating four models on three reasoning benchmarks, and finding three findings that argue for budget-conditioned evaluation protocols.
Rodrigo Guedes de Souza, Alison R. Panisson· 1 citation
This work evaluates three dense and mixture-of-experts models on BBQ and BBQ-V under seven conditions spanning batching, quantization, benchmark reduction, and their combinations, and compares accuracy, bias severity and prevalence, reasoning quality, subgroup behavior, subset-membership stability, runtime, and measured GPU energy against a full-benchmark BF16 baseline.
Ahmed El kady, Aravind Narayanan, Rehana Noorani et al.· 0 citations
Evaluating explainability methods requires more than a single faithfulness proxy. We present a modular benchmarking framework centered on quantitative XAI quality metrics: fidelity, stability, sparsity, computational cost, and faithfulness gap, plus an explicit method for operating the framework end-to-end. On the UCI Adult benchmark, we use a staged evidence protocol: a calibration/reproducibility stage (EXP1) followed by a primary com-parative/robustness benchmark (EXP2); the current merged recovery snapshot contains 299 committed result artifacts (99.7% artifact coverage) plus a 30-row SHAP recovery batch, yielding 275 analyzable unique runs out of 300 planned cells (91.7%). Across complete model-size blocks (5 models, N ∈ {50, 100, 200}), Friedman tests indicate significant method differences for fidelity (χ2 = 42.12, p = 3.78 × 10−9), stability (χ2 = 40.68, p = 7.65 × 10−9), sparsity (χ2 = 35.64, p = 8.92 × 10−8), faithfulness gap (χ2 = 45.00, p = 9.25 × 10−10), and runtime (χ2 = 30.44, p = 1.12 × 10−6). SHAP leads on fidelity/stability, DiCE leads on sparsity, and LIME remains fastest overall. We release the framework, operation protocol, and artifacts with explicit data-quality caveats for reproducible benchmark use under a quantitative-only claim scope.
Jonathan Herrera Vasquez, Miguel Herrero Uceda· Revista de investigación mul...· 0 citations
Claims of large lifts in A/B tests are widespread, yet many are only supported by small online experiments that are likely underpowered. Trustworthy A/B Patterns is a community replication effort to evaluate selected patterns at high statistical power. We report on eight A/B tests across four patterns (rounded buttons, page performance, coupon-code field, and sticky call-to-action), with a median of 2.4M users per experiment and 80% power at our pre-selected minimum detectable effects (MDEs) of 0.3% to 2.2%. We find that: (1) even at this scale, we did not have enough power for key business metrics, such as revenue per user and purchase conversion rate, within practical time horizons, and had to resort to surrogate metrics, such as click-through rate, add-to-cart rate, and capped add-to-carts (count); (2) previously reported effects for these patterns are highly exaggerated: across all eight replications, estimated effects were substantially smaller than previously claimed; only two showed statistically significant effects in the expected direction at α=0.05, and one was statistically significant in the opposite direction. While it is possible that some patterns have larger effects in certain conditions, we believe it is more likely that many of the prior estimates came from underpowered experiments (power below 50%), which exaggerate treatment effects. We conclude with lessons from running the community project for over one and a half years.
Ron Kohavi, Jakub Linowski, Lukas Vermeer et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.