Large Language Model Architectures and Their Trade-offs in Efficiency and Understanding
Abstract
As an important trend in the development of artificial intelligence, the large language model (LLM) is committed to building a two-way human-computer interaction, which has excellent performance in dialogue, real-time feedback, task execution, and so on. Research LLM architectures and their trade-offs in efficiency and understanding. Based on the references, this paper summarizes the basic theoretical framework of LLM technology, including its definition, key features, development history, and core technologies. Then, the existing literature is quantitatively analyzed, and the research hotspots of LLM technology are analyzed by using Citespace bibliometric tools. Based on the LLM, this paper mainly classifies its technical architecture, which is divided into pure decoder, encoder-decoder, and sparse hybrid expert. To reflect its interactive ability, this paper supplements it from two aspects: dialogue depth and multimodal support, and makes a comprehensive comparison of several existing mainstream LLMs. LLM can deal with complex decision problems, and is an important support for intelligent decision technology by facing the human-computer interaction mechanism to realize dynamic adjustment and self-optimization.