SALR is proposed, a schema-anchored latent reasoning method for LF construction that performs multi-step reasoning by generating continuous thoughts in the model's hidden states, thereby delaying the explicit commitment to LF decisions.
Guang-Ze Gao, Zi-Xuan Li, Si-Kui Zhang et al.· 0 citations
This paper proposes a scalable framework for image restoration and enhancement built upon the reinforcement learning (RL) paradigm, and confirms the model’s superior few-shot and zero-shot capabilities compared to existing methods, as well as its flexibility in addressing multiple competing objectives.
Juan Wang, Ke Zhang, Chun-Feng Yuan et al.· International Journal of Com...· 0 citations
SPAR is proposed, a novel joint semantic-geometric encoding architecture that explicitly isolates transient dynamic noise prior to latent space aggregation that structurally couples motion estimation with multi-view visual and semantic learning and reveals a strong inter-task synergy between photometric scene reconstru...
MMAgent-R$^2$, an agentic mRAG framework that integrates visual reranking and active rejection as its internal verification mechanism, is proposed and achieves joint optimization of external retrieval, internal verification, and answer generation via GRPO training.
Tao Zhang, Ziqi Zhang, Zongyang Ma et al.· arXiv.org· 0 citations
Gaussian Mixture Modeling for Event-Aware Visual Allocation is proposed, which leverages Gaussian Mixture Models to model event-level structure from discrete frame-wise observations and achieves comparable performance to baseline selection methods while utilizing only approximately half of the visual token budget.
Yifan Lu, Ziqi Zhang, C. Yuan et al.· arXiv.org· 0 citations
A Semantic-Retrieval-Augmented Detector (SRA-Det) is proposed that uses an attention-based module to retrieve multiple semantic facets from token-level text features, and a soft-min matching rule that behaves like a differentiable logical AND over these facets, ensuring that all key attributes are satisfied.
Li Yang, Boyu Cai, Wei Liu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.