Review
Jul 2026
FabriVLA: A Lightweight Vision-Language-Action Model with Conformal Action Chunk Uncertainty
FabriVLA is presented, a lightweight VLA that fuses shallow and intermediate VLM layers to preserve fine-grained visual features, and gates self-attention among action tokens so that its flow matching head admits inter step structure only as far as training warrants.
Shiyuan Yang, Borong Zhang, Jizheng Zhang et al.
· 0 citations