Application of machine learning and data analysis in construction project cost disputes: automated processing and risk control
Abstract
Addressing the challenges of processing multi-source heterogeneous data and the lagging identification of cost dispute risks in construction project settlement auditing, this paper constructs a machine learning-based numerical-semantic dual-path cost dispute precise identification model. This study extracts key features from three dimensions—numerical sensitivity, semantic conflict, and project environmental background—to construct high-dimensional feature vectors. The core architecture of the model consists of parallel dual paths: Path A utilizes the XGBoost algorithm to capture explicit numerical risks within engineering quantity and price deviations; Path B leverages an Attention-BiLSTM model to mine implicit fingerprints of rights and responsibilities conflicts within contract and change order texts. By introducing a gated fusion unit, the system achieves non-linear mapping and trade-off between numerical probabilities and semantic indices, and performs global parameter tuning in conjunction with a Bayesian optimization strategy. Case study results confirm that the model performs excellently on a dataset of 150 real settlement nodes, achieving an F1-score of 0.901 and an accuracy of 92.5%, with identification performance significantly surpassing traditional single-path identification models. Practical evaluation data indicates that the model can shorten the time consumed for individual audits by 65% and provide early warnings on average 14 d ahead. This study provides high-precision technical means for the intelligent auditing of construction project costs, offering theoretical support and practical reference for achieving the transformation from traditional experience-based auditing to data-driven risk prevention and control.