Skip to content

Author

Yunfei Chu

We have 6 of 25 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning

Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities independently, leaving the capability of audio-visual joint reasoning poorly evaluated and insufficiently elicited. We addre...

Junming Lin, Yuxuan Wang, Zhen-Xin Lei et al. · 0 citations
Preprint Aug 2026

LongAudioSpan: Spanning the Duration and Depth of Audio Comprehension

General audio comprehension now covers speech, sound, and music over durations from seconds to hours, driven by large audio-language models (LALMs) that are increasingly omni-modal. Yet the benchmarks that test them still rely on clips of seconds, where scores saturate and models converge; recent long-form efforts exte...

Wen Huang, Yunfei Chu, Meng Gao et al. · 2 citations · ⚡1
#artificial intelligence Preprint Sep 2026

OmniEcho: Audio-Visual Spatial Understanding for Omni-Modal Embodied Agents

Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents. In particular, it is still unclear how to effectively evaluate and model spatial audio understanding in embodied settings. To address this gap, we introduce...

Rui-Xun Liu, Yuxuan Wang, Jia-Cheng Xie et al. · 1 citation · ⚡1
#computer vision Preprint Sep 2026

Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental qu...

Qi Chen, Yunfei Chu, Haolin He et al. · 1 citation · ⚡1
#artificial intelligence Preprint Sep 2026

OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, without a separate text question, external c...

Haolin He, Yunfei Chu, Qi Chen et al. · 0 citations
Preprint Aug 2026

AudioSpan: Spanning the Duration and Depth of Audio Comprehension

This work introduces AudioSpan, a benchmark that spans both duration and depth: it pairs audio from 10 minutes to over 2 hours with 3,240 questions across three cognitive levels, namely perception, understanding, and reasoning.

Wen Huang, Yunfei Chu, Meng Gao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.