Skip to content

Author

Yutong Song

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Jul 2026

The empowerment of Science of Science by large language models: New tools and methods

Large language models (LLMs) have exhibited exceptional capabilities in natural language understanding and generation, image recognition, and multimodal tasks, charting a course toward artificial general intelligence and emerging as a central issue in the global technological race. This article conducts a comprehensive review of the core technologies that support LLMs from a user’s standpoint, including prompt engineering, knowledge-enhanced retrieval-augmented generation (RAG), fine-tuning, pre-training, and tool learning. In addition, it traces the historical development of Science of Science (SciSci) and presents a forward-looking perspective on the potential applications of LLMs within the scientometric domain. Furthermore, it discusses the prospect of an AI agent-based model for scientific evaluation and presents new research fronts in detection and knowledge graph building methods with LLMs.

Guoqiang Liang, Jingqian Gong, Mengxuan Li et al. · 0 citations
Preprint Aug 2026

AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models

Audio clustering is a fundamental task for organizing rapidly growing speech collections, supporting applications such as conversational analysis and speech-driven discovery. However, existing methods rely on fixed acoustic similarity metrics or ASR-based text pipelines, limiting their ability to reorganize the same audio collection under different user-specified perspectives, especially when clustering depends on both linguistic and paralinguistic cues. We introduce audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments. To study this setting, we construct AudioLens-Bench, a benchmark spanning multiple application domains and evaluating both in-perspective and cross-perspective generalization. We further propose AudioLens-R1, an end-to-end large audio-language model trained with reasoning distillation and preference optimization. Experiments show that AudioLens-R1 consistently outperforms all baselines, improving overall ARI by 12.99 points and V-measure by 11.62 points. These results demonstrate the promise of native audio-language models for flexible, perspective-conditioned structure discovery over speech collections.

Wenjun Huang, Qiao-Song Chu, Tiger Shao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.