A search-based metamorphic testing approach that identifies a minimal set of transformations on underwater images to induce incorrect model predictions, thereby revealing VLM failures and derive lessons for software engineering practitioners and researchers working on quality assurance of VLM-based software systems.
Muhammad Yousaf, Aitor Arrieta, Shaukat Ali et al.· 0 citations
To ensure the overall quality of AI-enabled software, not only traditional software components but also AI components need to be tested and repaired. Among AI components, Transformer models are increasingly integrated into software systems, which makes their misbehaviors critical. Although prior work in the software engineering community has proposed deep neural network (DNN) repair methods, most overlook Transformer-specific structures. We propose RepTran, a search-based repair method for Transformer models. It targets their feed-forward networks (FFNs), which play a central role in the architecture. RepTran identifies suspicious weights by combining two types of scores: a variance-based neuron score and an existing bidirectional score. It then iteratively optimizes these weights using differential evolution. Our evaluation includes 18 fault benchmarks constructed from CIFAR-100 and Tiny-ImageNet. We compare RepTran against three baselines: random weight selection, Arachne (a state-of-the-art DNN repair method), and ArachneW, which enables Arachne to control the number of selected weights. RepTran achieved an average repair rate of 74.7%, statistically outperforming random selection and Arachne across all benchmarks. Effect size analysis revealed that RepTran achieved higher repair rates than ArachneW regardless of the number of selected weights. These results suggest that RepTran is effective for enhancing the reliability of AI-enabled software.
Yuta Ishimoto, Paolo Arcaini, Fuyuki Ishikawa et al.· arXiv.org· 0 citations
This work creates a novel open-source symbolic traffic model EvoDrive designed specifically for EC research, which outputs LLM-readable snapshots and shows that LLMs + EvoDrive with CLEAR can reduce error by more than 20% compared to without CLEAR, with statistically significant results.
Peter J. Bentley, S. Lim, Fuyuki Ishikawa et al.· Annual Conference on Genetic...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.