This work explored powerful explainers which use rules, where explanation instability stemming from training data becomes more apparent, and found that the rule-based method employed, BARBE, sharply increased in fidelity and stability when trained with the modified process, making BARBE+PBC which exceeded other methods that improve stability like S-LIME and LORE.
This work presents a comprehensive literature review of the current state of the subfield of XAI that consist of causality-motivated post-hoc XAI methods, and a causal framework for categorising XAI is introduced, and three types of post-hoc XAI methods are identified: observational methods, internally causal methods and externally causal methods.
Anna Rodum Bjøru, Helge Langseth, Inga Strümke et al.· Machine-mediated learning· 1 citation
Explaining machine-learning models is increasingly important for decision-making and consumer trust, yet it is widely believed to come at a cost: existing Explainable AI (XAI) methods suffer from a persistent accuracy-explainability trade-off. We argue that this trade-off is not fundamental, but an artifact of treating explanation and prediction as separate objectives; when properly coupled, they become complementary, so that equipping a model to explain itself improves, rather than degrades, its accuracy. We introduce the Rashomon Explanation paradigm, which builds a set of faithful, prediction-guiding explanations rather than a single one, and prove that this set is generally non-empty and that explanation fidelity bounds the performance of the models it guides. To explore this set, we propose RashomonLLM, an Explanation-Prediction-Reflection agentic workflow that generates explanations in natural language by iteratively aligning them with predictions, and we prove it converges and recovers the full set. Across customer-churn classification, clinical survival regression, and industrial click-through prediction on large-scale live-streaming logs, RashomonLLM significantly outperforms state-of-the-art prediction and XAI baselines on both accuracy and explanation quality, with gains driven by explanation fidelity and robust to distribution shifts, temporal splits, and seeds. Our framework thus advances business performance while laying the groundwork for consumer trust.
This work shows that pooled refitting recomputes the inverse-Gram geometry used to weight source evidence, which can reverse shared preferences, and derive exact and approximate preservation conditions, and develops a three-stage audit that traces strict pairwise reversals through decision changes to task-defined utility loss.
The nature of test-time exploration in RLVR-trained LLMs is investigated by employing controlled maze-solving experiments and extracting a tree structure from mathematical reasoning traces (BODHI-Trees) based on semantic equivalence to delineate between entropy arising from stylistic variations and genuine inferential branching.
Soumadeep Saha, Krish Sharma, Akshay Chaturvedi et al.· 1 citation
Statistical tests are often asked to do too much. A single reported result is expected to describe what the observed data say, reassure readers about repeated-sampling behavior, and remain convincing when the working model is perturbed. Those tasks are connected, but they are not equivalent. Fisherian inductive inference and Neyman-Pearson decision theory clarify the first two; robust testing, sensitivity analysis, fragility measures, multiverse analysis, and distributional-stability methods speak to the third. I propose Evidence-Calibration-Stability (ECS) as a framework for keeping these roles separate while reporting them together. Evidence is post-data. Calibration belongs to the design or procedure. Stability is the post-data distance from the benchmark analysis to a conclusion-reversing perturbation within a declared model neighborhood. Full ECS support is conjunctive: a strong coordinate cannot rescue a failed one. For finite-dimensional affine perturbations, I derive an exact ellipsoidal stability radius. For smooth nonlinear margins, a uniform quadratic-remainder condition yields a certified lower bound over a declared neighborhood, showing when the affine formula is only a surrogate. I also establish coordinate invariance and a matrix extension for multiple claims, and distinguish confirmatory calibration from descriptive calibration profiles when prespecification is unavailable. Simulations for the one-sample t test and Student's historical sleep data show that the three coordinates can lead to different interpretations. ECS is a formal synthesis, not a claim that evidence, power, or robustness is itself new.
Subir Hait· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.