Skip to content
Conference

CoLSM: Collaborative Large and Small Models for Automatic Software Generation

Aug 2026 · 2026 12th International Conference on Big Data and Information Analytics (BigDIA) · pp. 35-42 · 0 citations · 32 references

Abstract

Large language models (LLMs) support human-in-the-loop code development by rapidly generating high-quality code snippets. However, they still face prominent challenges in fast and efficient deployment on edge environments. Such challenges mainly involve heavy computation costs, poor domain accuracy, unbalanced collaboration efficiency and inconsistent cross-model knowledge. This study proposes CoLSM, a new collaboration mechanism guided by mixture experts for automatic software generation. It establishes a hierarchical and iterative working pipeline. A mixture-expert router assigns tasks dynamically. Large models take charge of system architecture and complex logic design. Domain-adapted small models refine code details, optimize resource usage and ensure security compliance. This mechanism integrates an abstract syntax tree based synchronization module to resolve cross-model conflicts and embeds a quality feedback loop to support adaptive iterative optimization. We evaluate the proposed CoLSM on a self-built multi-scenario software generation dataset. Experimental results demonstrate that CoLSM improves software generation accuracy and functional consistency by 4.3% and 5.7%, respectively. It also reduces inference latency by 20.3% and energy consumption by 14.8%. CoLSM effectively combines the respective advantages of large and small models. It realizes accurate and low-cost automatic software generation and provides reliable technical support for agile automated software development.

View source

Similar papers

Aug 2026

Cross-Model Collaboration for Enhancing LLM-Based Code Generation

Findings indicate that cross-model collaboration offers a practical and parameter-efficient alternative to scaling up monolithic models for code generation and maintains competitive accuracy when only 20% of test cases are available for diagnostic feedback.

Jiangping Huang, Wen-Guang Ye, Weisong Sun et al. · 0 citations
Book Open access Aug 2026

Towards Lightweight Domain Model Reverse Engineering from Source Code

This paper proposes an automated approach to extract domain models from source code using lightweight, locally deployable LLMs and achieves high F1-scores on a dataset of ten projects, each comprising a curated domain model and its corresponding implementation, while remaining fully executable on locally deployable LLM...

Kévin Delcourt, Meriem Ben Chaaben, Abdelhamid Rouatbi et al. · 1 citation

Hunk-Constrained DPO: Segment-Level Optimization for Secure and Correct LLM Code Generation

Hunk-Constrained Direct Preference Optimization is introduced, a training framework that unifies security hardening and functional correction in large language models and demonstrates that HPO achieves substantial security improvements—up to 28 percentage points—while preserving or enhancing functional correctness.

Qian-Shuo Huang, Xin Yin, Xin-Rui Li et al. · 0 citations
Preprint Aug 2026

DeepRepro: State-Aware Subplanning for Paper-to-Code Reproduction in Evolving Repositories

DeepRepro dynamically transforms evolving repository states and runtime feedback into fine-grained implementation subplans, keeping planning aligned with execution throughout repository construction, and consistently outperforms strong scientific and commercial code-agent baselines.

Hong-Ru Song, Ru-Qing Zhang, Jia-Feng Guo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Retrieval-Augmented Generation for Scientific Code Understanding

Large language models have become central to modern coding assistants, but state-of-the-art systems such as Claude Code or Codex rely on very large, cloud-hosted models with significant computational cost and data-privacy implications. This work investigates whether a useful, fully local coding agent can be built aroun...

Aaron Nobile, Andreas Adelmann, Mohsen Sadr · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.