UvA-DARE (Digital Academic Repository) TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free Tuning
This paper proposes TWIST, a twin-expert stepwise tuning module that modifies the decoder of the language model using one frozen module pre-trained on image understanding tasks and another learnable one for visual grounding tasks, which allows the MLLM to retain previously learned knowledge and skills, while acquiring what is missing.