LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models
This work introduces LIBERO-Para, a controlled benchmark that independently varies action expressions and object references for fine-grained analysis of linguistic generalization in VLA models, and proposes PRIDE, a metric that quantifies paraphrase difficulty using semantic and syntactic factors.