This claim for multi-turn, tool-calling agents, where it now matters most, is tested for post-training quantization to 4-bit weights and diagnostics, the per-channel error rate and success under a shrinking budget come from logs benchmarks already collect.
On HealMed, performance declined most in low-resource languages, although the size of the gap varied markedly across languages and models, whereas many open-source and medically specialized models showed larger and less consistent gaps.
Yingjian Chen, Fan Gao, Sherry T. Tong et al.· 0 citations
Funnel of Thoughts (FoT) is introduced, an inference-time method that preserves the full 32-trajectory voted accuracy while halving its attention FLOPs, a 28.8% reduction in full-model inference cost.
Chanhee Park, Sun Han, Jeongho Yoon et al.· 0 citations
This paper proposes Cross-lingual Ranking Preference Optimization~ (CRPO), a novel framework that leverages robust preference knowledge from English to facilitate preference alignment in the target language, thereby enhancing language adaptation and output quality.
Seungyoon Lee, Minhyuk Kim, Jungseob Lee et al.· 0 citations
This work releases LAMAR, a language aware multilingual cross encoder trained to account for both semantic relevance and language coherence, which achieves the best performance overall and across all languages examined individually on general multilingual reranking benchmarks.
Seongtae Hong, Youngjoon Jang, Jungseob Lee et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.