Skip to content
Open access

A Named Entity Recognition Method for GIS Defect Texts Incorporating an Engineering Format-Aware Masking Strategy

Aug 2026 · Energies · Vol 19, pp. 4004 · 0 citations · 18 references

TL;DR

An engineering format-aware masking strategy is proposed that improves entity recognition for phase-related structures, engineering abbreviations and equipment hierarchy fragments and achieves precision, recall and F1 scores of 0.90, respectively.

Abstract

Named entity recognition (NER) is a key technique for extracting entities such as equipment, components and defect types from GIS defect texts, providing a basis for subsequent knowledge graph construction. However, GIS defect texts contain many engineering structures, including engineering abbreviations, equipment numbers and phase identifiers, making it difficult for general-purpose models to stably recognize their semantic associations and entity boundaries. To address this problem, this paper proposes an engineering format-aware masking strategy. The strategy identifies candidate fragments using format rules for phase identifiers, measurement value-unit patterns and engineering abbreviations and preferentially selects them as perturbation targets to strengthen the model’s understanding of engineering structures and their contextual relationships. Bidirectional long short-term memory is used to extract bidirectional sequence features, and a conditional random field is used to model transition constraints between labels and obtain the globally optimal label sequence. The results show that the proposed model achieves precision, recall and F1 scores of 0.89, 0.92 and 0.90, respectively. The analysis indicates that the proposed method improves entity recognition for phase-related structures, engineering abbreviations and equipment hierarchy fragments.

Read PDF

Similar papers

Aug 2026

A large language model-based question-answering system for crack information

This research provides a highly accurate, scalable, and reliable framework for automated bridge defect analysis, offering a practical methodology to enhance data utilization in bridge management.

Lu-yang Zhang, Xuzhao Lu, Fengzong Gong et al. · 0 citations
Open access Aug 2026

Automatic Extraction of Key Information from Financial Statements Using a Multi-Task Learning Model

Financial statement information extraction and technical document data management both face significant challenges due to complex formats, scattered layouts, and intricate proximity-based semantic relations within semi-structured documents. These challenges are particularly evident in engineering enterprises involving electromagnetic devices, antenna systems, radio-frequency equipment, and related technical service activities, where accurate extraction of financial and operational information is required for reliable reporting and performance analysis. This paper proposes an automatic extraction method based on multi-task learning. The model employs a shared encoder, in which text features and layout features are fused to process four related tasks simultaneously: financial named entity recognition, key-value extraction, entity-value relationship classification, and report item classification with DistilRoBERTa. Experiments based on 500 annual reports show that the accuracy of entity recognition reaches 96.8%, the F1 score of value extraction reaches 88.7%, and the F1 score of relationship classification reaches 91.5%. The model performs particularly well in structured sections such as the balance sheet and income statement. Although its performance declines to some extent in the most complex and unstructured sections of financial statements, such as notes, the overall results demonstrate its effectiveness in improving the accuracy and semantic consistency of financial disclosure information extraction. This study provides a technical reference for intelligent document understanding and reliable data extraction in financial reporting scenarios involving engineering-oriented enterprises and electromagnetic application industries.

C. Chen · 0 citations
Open access Aug 2026

Domain-Specific Retrieval-Augmented Generation for Metallurgical R&D Knowledge Bases: A Hybrid Graph-Enhanced Approach

This work evaluates a confidence-adaptive graph-enhanced retrieval layer for metallurgical RAG in a controlled synthetic benchmark with explicitly specified generation and evaluation rules and shows how the retrieval rule behaves under controlled conditions.

Viktor A. Vedeneev, V. Kondratiev, A. Nazarychev et al. · 0 citations
Preprint Aug 2026

Domain-Specific Text Embedding Models for Entity Resolution

General-purpose text embedding models are designed to capture semantic similarity but are not optimised for distinguishing entity records that represent the same real-world business or person. This limitation affects applications such as entity resolution and duplicate record retrieval, where small textual differences may either preserve or change identity. This paper investigates whether domain-specific triplet fine-tuning can adapt pretrained embedding models for identity-sensitive retrieval. A synthetic dataset of business and person records was created with identity-preserving variations and challenging non-matching examples. Two widely used embedding models were evaluated before and after fine-tuning using a margin-based similarity evaluation. The results show substantial improvements in separating true matches from highly similar non-matches, demonstrating that domain-specific triplet training can effectively reshape general-purpose embedding spaces for entity retrieval. These findings suggest that targeted fine-tuning provides a practical approach for improving embedding models in data quality management and information retrieval applications.

Khajesh Sapram, S. Raju, Kishore Konda · 0 citations
Book Open access Aug 2026

Towards Automated P&ID Digitization: Graph-Based OCR Consolidation and Global Symbol-Tag Association

Piping and Instrumentation Diagrams (P&IDs) are essential engineering documents, but many remain available only as scanned PDFs, limiting their integration into digital workflows. Automatic extraction of structured information from these drawings is challenging due to large document sizes, small text annotations, and ambiguous symbol-tag relationships. This paper presents an end-to-end framework for automatic instrument-tag extraction and association from scanned P&IDs. A tiled OCR strategy with graph-based text merging improves text completeness and reconstructs fragmented engineering tags. For symbol detection, an RF-DETR model fine-tuned on the Dataset-P&ID benchmark achieves 99.96% mAP@50 and 99.97% precision, while SAHI-based inference slicing improves performance on large drawings. To automate symbol-tag association, we formulate the problem as a minimum-cost bipartite matching task that combines geometric, semantic, and spatial cues and solves it globally using the Hungarian algorithm. Results on Dataset-P&ID demonstrate accurate tag reconstruction, highly reliable symbol-tag associations, and a substantial reduction in manual annotation effort, providing an effective foundation for large-scale P&ID digitization. The code is available at https://github.com/dimitri009/STA.

Nguinwa Mbakop Dimitri Romaric, Simone Marinai · 0 citations
Aug 2026

Knowledge-Graph-Supported Indicator Generation for HSE Management System Evaluation Using BERT-BiLSTM-CRF

To improve the systematicity and traceability of HSE system evaluation indicators, this study addresses limitations in traditional indicator selection, including strong dependence on expert experience, high subjectivity, overlapping indicator boundaries and unclear source evidence. Based on a first-level indicator framework determined by enterprise experts, HSE standards, system documents and related literature were used as the corpus. A domain-specific entity annotation scheme was designed, covering management measures, responsible actors, evidence information, risk objects, resource support, tools and technologies, and improvement actions. A BIO-annotated dataset was then constructed, and the BERT-BiLSTM-CRF model was employed for domain entity recognition. On this basis, second-level evaluation indicators were generated and screened through knowledge fusion, graph-path retrieval and topic consolidation. The results show that the BERT-BiLSTM-CRF model achieved a precision of 91.2%, a recall of 89.8% and an F1-score of 90.5%. Based on entity recognition and topic consolidation, an HSE system evaluation indicator system comprising 9 first-level indicators and 36 second-level indicators was developed, providing a basis for subsequent indicator weighting and comprehensive evaluation modeling.

Kexin Sun, Hui-Ling Na, Jianwei Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.