Skip to content

Author

Jordi Escriu

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Towards AI-assisted metadata generation for improved description of geospatial data

Metadata are fundamental components of spatial data infrastructures, enabling the discovery, evaluation, and reuse of datasets. Their creation and maintenance, typically performed manually, can be costly, time-consuming, and prone to inconsistencies. This work investigates the feasibility of using Generative AI (GenAI) to assist human operators working for data producers in creating structured dataset descriptions compliant with GeoDCAT-AP standards. We address the following research question: to what extent can layered prompt engineering strategies improve the quality of metadata descriptions generated through Large Language Models (LLMs), and how do individual prompt components-role definition, content rules, template structure, and few-shot examples-interact with model selection to affect structural compliance, semantic similarity, and factual reliability? We propose a six-layer prompting framework and evaluate seven ablation strategies, each selectively disabling specific layers, using two LLMs (qwen3-32b and qwen3-coder-30b-a3b-instruct) and eight geospatial datasets. Generated descriptions are assessed using eight automated metrics spanning lexical quality, structural compliance, and semantic similarity, complemented by targeted expert evaluation of factual accuracy. Results reveal a clear hierarchy of layer impact. The input template, used as the reference structure of the description, is the most influential component, driving both structural compliance and readability. Few-shot examples are the second most impactful layer, substantially reducing redundancy and improving semantic alignment. Content rules do not measurably improve surface-level output quality, but serve as critical safeguards, encouraging models to flag missing information rather than fabricate content. Expert role definition contributes the least measurable effect. Both LLMs produce semantically comparable outputs under full guidance but diverge without structural constraints. We recommend the full prompting configuration for production use and provide practical guidelines for balancing prompt complexity, output quality, and factual reliability in LLM-assisted metadata generation workflows.

M. Di Leo, Ilyas Tiouassiouine-Maes, Jordi Escriu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.