It is asserted that the present contribution consists of an interpretable domain palette, a constructed benchmark of diverse tabular datasets, and reproducible code and data to enable further research on domain discovery and domain-aware tooling for tabular data.
AssistEM, a framework for efficient LLM adaptation to EM via principled data selection, demonstrates that selective fine-tuning not only accelerates adaptation but also improves training efficiency (requiring fewer GPU hours), enabling open-source LLMs to rival–and in some cases outperform–closed-source models.
John Bosco Mugeni, Steven J. Lynden, Toshiyuki Amagasa et al.· International Journal of Dat...· 0 citations
This research focuses on enterprise profiling in scenarios where large volumes of diverse texts—such as registration records, annual reports, news articles, and bidding notices—are continuously generated. Instead of relying solely on a single data representation or classification model, we developed a comprehensive natural language processing (NLP) pipeline for extracting key information and identifying industries. The pipeline consists of several steps. First, we use a BERT-BiLSTM-CRF model to identify core e/nterprise entities. Then, we combine TF-IDF with BERT embeddings to create a hybrid feature scheme that captures both lexical cues and contextual semantics. To address the challenge of imbalanced industry labels, we apply SMOTE in the dense semantic space and pair it with Focal Loss to enhance learning for minority classes. Additionally, we introduce a Stacking strategy to integrate outputs from different models, making predictions more stable. Tests on a self-compiled dataset covering ten national economic sectors and about 50,000 enterprises show that our method achieves a macro-F1 score of 95.4%. It outperforms traditional machine learning baselines and single deep learning models, offering more reliable recognition for minority classes. These results suggest that our framework is well-suited for applications such as supply chain partner discovery, industrial mapping, and targeted investment promotion.
Xin-Yi Xu· Applied and Computational En...· 0 citations
A dataset and annotation tool to support the development of German-language Tabular Question Answering systems, with a specific focus on sustainability-related information, exclusive use of the German language and a strong emphasis on information retrieval from tables embedded in sustainability reports.
Introduction Large organisations maintain heterogeneous datasets in data lakes, where schema variability, inconsistent formats, and semantic ambiguity complicate reconciliation. Achieving a unified view of entities requires methods that capture both structural equality and semantic relationships. Methods We propose a transformer-based methodology that adapts pre-trained language models (PLMs) to tabular data by generating metadata-enriched embeddings (table title, column name, type, and statistics). These embeddings are compared using mutual top-K similarity, value-level verification, and Facebook AI Similarity Search (FAISS) for discrepancy detection. Ground-truth labels were established through manual annotation of 1,000 column pairs per dataset, with three annotators achieving substantial agreement (Cohen’s κ = 0.82). Results Experiments on large-scale tabular data (185,909 tables) demonstrate that semantic embeddings efficiently uncover relationships and discrepancies, achieving precision of 0.958 at τ = 0.9 and F1-scores ranging from 0.77–0.87. Benchmarking against baselines (exact matching, Jaccard similarity, edit distance, TF-IDF/BM25, sentence-transformer embeddings, DeepJoin-style embeddings, and schema-name-only matching) shows that lexical baselines achieve high precision but poor recall, while semantic baselines capture relationships but underperform on heterogeneous tables. In contrast, our metadata-enriched transformer embeddings consistently achieved the highest F1-scores. Discussion Unlike prior schema-aligned or query-driven approaches such as DeepJoin, WarpGate, and Lotus, this study introduces a unified reconciliation pipeline that integrates equality-based and semantic matching for large-scale tabular data. The key contribution is a scalable and generalisable reconciliation methodology that operationalises transformer architectures for automated discrepancy detection in heterogeneous environments, establishing both methodological novelty and practical effectiveness.
Carl Du Plessis, Mike Wa Nkongolo· Frontiers in Artificial Inte...· 0 citations
This study explores a semantic variation methodology to augment training data by generating question-answer pairs with explicit control over semantic similarity, and shows that semantically controlled augmentation improves domain-specific knowledge acquisition while preserving consistency.
Alexander Chen, Caroline Tang, Jennifer Sleeman· TEXT2KG/BiKE@ESWC· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.