Skip to content
Preprint

Reasoning-Oriented Post-Training and Inference-Time LoRA Rescaling for Audio-Dependent Question Answering

Aug 2026 · 0 citations · 20 references
Computer Science

TL;DR

A structured Chain-of-Thought framework is introduced that decomposes the reasoning process into question analysis, question type, audio evidence, and reasoning, and how task-specific LoRA adaptation affects the two backbones is analyzed and inference-time rescaling of trained LoRA adapters is explored.

Abstract

Audio-Dependent Question Answering (ADQA) requires Large Audio-Language Models (LALMs) to answer questions whose correct answers depend on the given audio content. Successful ADQA requires accurate audio perception, identification of question-relevant evidence, and cross-modal reasoning. Using the official ADQA dataset of DCASE 2026 Task 5, we investigate reasoning-oriented post-training with Low-Rank Adaptation (LoRA) and inference-time LoRA rescaling for both Qwen2.5-Omni and MOSS-Audio-8B-Thinking. We introduce a structured Chain-of-Thought (CoT) framework that decomposes the reasoning process into question analysis, question type, audio evidence, and reasoning. We then analyze how task-specific LoRA adaptation affects the two backbones and further explore inference-time rescaling of trained LoRA adapters. Experiments on the development set reveal markedly backbone-dependent behavior: post-training improves the Qwen-based systems but substantially degrades MOSS-Audio under our supervised fine-tuning configuration. Moderate LoRA rescaling further improves the best Qwen system's top-1 accuracy from 58.93% to 61.05% and partially restores the performance of the fine-tuned MOSS-Audio models, while the best MOSS-Audio system achieves 67.70% top-1 accuracy. Our submitted systems ranked third overall and second among lightweight systems under 10B parameters in the challenge.

View source

Similar papers

Review Jul 2026

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering

DCASE~2026 Task~5 introduces Audio-Dependent Question Answering (ADQA), which tests whether large audio-language models answer from the audio rather than from textual priors. An Audio-Dependency Filtering (ADF) pipeline combines silent-audio probing, per-option perplexity, a large language model (LLM) commonsense check...

Haolin He, Renhe Sun, Zheqi Dai et al. · 1 citation
Open access 2026

Enhancing Auditory Reasoning in Large Audio–Language Models Via Supervised Fine-Tuning and Reinforcement Learning With Verifiable Rewards

The objective of this paper is to improve and analyze auditory reasoning in large audio–language models for audio question answering (AQA), where a model must infer the correct answer from acoustic evidence and textual answer options. Although reinforcement learning (RL) with verifiable rewards has recently improved re...

Jian Wang, Sheng-Yan Hao · 0 citations
#natural language process... Preprint Sep 2026

SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ

This paper describes our system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge. We adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline. A language model converts timestamped ASR into coherent event spans, which are expanded by a...

H. Le, L. Nguyen, Minh Tri Dao · 1 citation
Jul 2026

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning

Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or provide only coarse semantic signals. T...

Siqian Tong, Xuan Li, Chaozhuo Li et al. · 0 citations
Preprint Aug 2026

AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models

This work introduces audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments.

Wen-Jun Huang, Qiao-Song Chu, Tiger Shao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.