Multi-modal Boundary Testing of Vision–Language Models
This work constructs a signal by force-decoding a fixed set of candidate answers and develops a framework that manipulates the image and question text separately, making the boundary searchable for any behaviour formulated as a choice between candidate answers.
R. Kaiser
· 0 citations