Skip to content

Author

Ruoxin Xiong

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Evaluating Large Language Models Using Construction Management Certification Exams: Implications for Construction Education and Professional Preparation

Large language models (LLMs) are increasingly used by students and professionals in the architecture, engineering, and construction (AEC) sector for learning and information retrieval. However, their reliability in performing construction management (CM) knowledge tasks remains insufficiently characterized. This study introduces CMExamSet, a benchmark data set consisting of 689 multiple-choice questions compiled from four professional CM certification programs. The data set covers core CM domains, including project and program management, safety management, cost control, scheduling, contract administration, and related professional knowledge areas. Four contemporary LLMs were evaluated using a standardized zero-shot protocol with five repeated runs per question. Performance was assessed using accuracy, response consistency, subject-area analysis, and structured error annotation. Mean accuracy ranged from 79.2% to 90.0% across the four examinations, with high overall agreement observed in repeated runs. Performance varied systematically by domain: higher accuracy was observed in information retrieval-based domain tasks, such as engineering concepts, construction geomatics, and sustainability, whereas lower accuracy was observed in operationally oriented domains such as bidding and estimating, time and schedule management, and cost control. When errors occurred, conceptual misunderstandings were the most frequently observed error type across domains, indicating that incorrect interpretation or application of domain principles remained a common source of failure. Performance improvements in newer models were generally modest across subject areas. These findings provide empirical evidence on the capabilities and limitations of LLMs in CM knowledge assessment and underscore the importance of structured verification, domain grounding, and instructional guidance when integrating LLM tools into construction education and professional preparation.

Ruoxin Xiong, Yanyu Wang, Suat Gunhan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.