A Model of Software Code Quality Analysis, Assessment, and Improvement System
Abstract
Existing code analysis systems address individual aspects of software quality (complexity, duplication, style) but do not combine multi-level analysis, metric aggregation, and learning-based inference within a single formal framework, nor do they close the loop between refactoring outcomes and assessment. This paper proposes a formal model of a code quality analysis, assessment, and improvement system organized as a sequential-hierarchical architecture of nine functional components. The model introduces an information kernel that accumulates structural, metric, anti-pattern, duplication, and dependency evidence into a context-complete state, and combines deterministic detection rules with an ensemble of regression models for integral quality inference. End-to-end behaviour is formalized as a composition of mappings with explicit linear, kernel-mediated, and feedback edges. The model was evaluated on three publicly available Java projects (approximately 410 thousand lines of code (KLOC) and 3,100 entities) against SonarQube 10.4 and PMD 7.0. The proposed model attained $\mathrm{F}1=0.83$ for anti-pattern detection (vs 0.74 for SonarQube and 0.69 for PMD), $\mathrm{R}^{2}=0.81$ for integral quality prediction, and Spearman $\rho=0.78$ between systemassigned recommendation priority and expert ranking. The results indicate that joint use of multi-level analysis, ensemblebased inference, and feedback from refactoring validation yields measurably better anti-pattern coverage and quality estimation than rule-based baselines.