Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an...
Yu Huang, Jun-Gang Li, Zhiyuan Wang et al.· 0 citations
InterTab is a structure-aware framework for CoT reasoning over table images that interleaves chain-of-thought with tool calls that crop structure-aligned table regions, and improves the average accuracy of its backbone from 68.28% to 73.17% and achieves the best average performance among all compared methods.
Hanqian Li, Si-Rui Huang, Chen Ling et al.· 0 citations
By unifying four hallucination dimensions with paired question design, KnowHal addresses an important gap in existing evaluation frameworks and enables a more comprehensive assessment of hallucinations in MLLMs.
Ruihan Li, Ji-Yang Tan, Kai-Lin Jiang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.