Large language models (LLMs) have exhibited exceptional capabilities in natural language understanding and generation, image recognition, and multimodal tasks, charting a course toward artificial general intelligence and emerging as a central issue in the global technological race. This article conducts a comprehensive review of the core technologies that support LLMs from a user’s standpoint, including prompt engineering, knowledge-enhanced retrieval-augmented generation (RAG), fine-tuning, pre-training, and tool learning. In addition, it traces the historical development of Science of Science (SciSci) and presents a forward-looking perspective on the potential applications of LLMs within the scientometric domain. Furthermore, it discusses the prospect of an AI agent-based model for scientific evaluation and presents new research fronts in detection and knowledge graph building methods with LLMs.
Guoqiang Liang, Jingqian Gong, Mengxuan Li et al.· Journal of information scien...· 0 citations
Audio clustering is a fundamental task for organizing rapidly growing speech collections, supporting applications such as conversational analysis and speech-driven discovery. However, existing methods rely on fixed acoustic similarity metrics or ASR-based text pipelines, limiting their ability to reorganize the same audio collection under different user-specified perspectives, especially when clustering depends on both linguistic and paralinguistic cues. We introduce audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments. To study this setting, we construct AudioLens-Bench, a benchmark spanning multiple application domains and evaluating both in-perspective and cross-perspective generalization. We further propose AudioLens-R1, an end-to-end large audio-language model trained with reasoning distillation and preference optimization. Experiments show that AudioLens-R1 consistently outperforms all baselines, improving overall ARI by 12.99 points and V-measure by 11.62 points. These results demonstrate the promise of native audio-language models for flexible, perspective-conditioned structure discovery over speech collections.
Wenjun Huang, Qiao-Song Chu, Tiger Shao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.