Preprint
Aug 2026
Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace
Reading a base Qwen3-Omni with a logit lens at the audio-token positions, it is found that the answer to a spoken question becomes legible - in words - in the model's middle layers, before it emits any token.
Jiajun Fan, Jing-Yuan Li, Prashanth Gurunath Shivakumar et al.
· 0 citations