Skip to content
Review Open access

Better Prompts, Better Usefulness: A Systematic Review and Experimental Evaluation of Structured Prompting Techniques in Large Language Models

Jul 2026 · Big Data and Cognitive Computing · Vol 10, pp. 224 · 0 citations · 87 references

TL;DR

Results indicate that structured prompting significantly increases perceived usefulness compared to baseline approaches, with the combination of example-based conditioning and explicit reasoning scaffolding yielding the highest evaluations.

Abstract

Large Language Models (LLMs) have rapidly become central components of cognitive computing systems and AI-assisted knowledge work. However, the effectiveness of LLM-generated outputs depends not only on the model’s capabilities but also on the structure of the prompts used to guide them. This study investigates how structured prompting techniques influence perceived output usefulness in business-oriented tasks. First, we conduct a systematic literature review following PRISMA guidelines to identify, classify, and synthesize existing prompt enhancement strategies. The review leads to the development of a taxonomy distinguishing task-alignment techniques (e.g., one-shot and few-shot prompting) from reasoning-transparency techniques (e.g., Chain-of-Thought prompting). Building on this taxonomy, we design a controlled experimental study in which knowledge workers evaluate LLM-generated outputs across analytical and summarization tasks. Using linear mixed-effects modeling, we assess the impact of prompting techniques and the moderating role of Generative AI usage frequency. Results indicate that structured prompting significantly increases perceived usefulness compared to baseline approaches, with the combination of example-based conditioning and explicit reasoning scaffolding yielding the highest evaluations. The moderating effect of usage frequency is not statistically significant, suggesting that the benefits of structured prompt design are robust across different experience levels. These findings position prompt structure as a practical cognitive interface mechanism and provide evidence-based guidelines for enhancing human–AI interaction in cognitive computing environments.

Read PDF

Similar papers

Open access Jul 2026

A Systemic Intervention for Human-Artificial Intelligence Co-Design of Lesson Plans: Integrating Pedagogical Theories into Prompt Engineering

Large language models (LLMs) can assist lesson planning, but simple prompts often yield incomplete and misaligned outputs. This study proposes a three-step prompt framework grounded in Bloom’s taxonomy, Adaptive Control of Thought-Rational (ACT-R) theory, Gagné’s nine events, and problem-chain theory, decomposing planning into objective, unit, and activity design. Using three DeepSeek models (R1, V3, 32B) and five prompting strategies, 150 lesson plans were generated on ten computer networking topics. Coverage of Gagné’s nine events and functional quality were evaluated via an LLM judge and human validation. All theory-based strategies significantly outperformed naive prompting, raising Gagné’s event coverage above 90% in the full corpus and from 74.1% to 89.8–93.5% in human ratings. Functional quality scores improved by up to 17.3% (LLM judge) and 53.6% (human raters). Gagné’s five-stage design outperformed ACT-R’s three-stage design under base conditions, while problem-chain guidance benefited ACT-R substantially. Model capability moderated gains: smaller models benefited most in structural completeness, stronger reasoners achieved higher absolute quality. These findings demonstrate that pedagogically grounded, multi-stage prompts are designed to reconfigure teacher-artificial intelligence (AI) interaction from passive output consumption toward structured collaborative design, offering a scalable intervention for integrating LLMs into instructional workflows.

Yinan Lu, Weinuo Li, Yuesheng Cai · 0 citations
Open access Jul 2026

Prompting for Independent Learning: An Evaluation of Tutoring Behaviors in GenAI

A simulation-based textual analysis of prompt design evaluates a frontier large language model as a tutor across 60 scripted sessions on a single topic and point to dynamic, dialogue-aware prompting alongside explicit SRL scaffolding.

Kendall Hartley, Fabiola Sáez-Delgado, Javier Mella-Norambuena · 0 citations
Review Open access Jul 2026

Automatic Question Generation with Large Language Models: A Survey

Test-based learning is effective in fostering knowledge retention, but manually creating assessment questions remains time-consuming and limits personalized student practice. The advent of Large Language Models (LLMs) has introduced new possibilities for Automatic Question Generation (AQG). Motivated by this context, this survey provides a comprehensive overview of AQG using LLMs, focusing on educational applications. Following the PRISMA methodology, we reviewed 132 studies published between 2023 and 2025. Our contributions include a taxonomy of question types by response openness, an analysis of AQG efforts across knowledge fields, educational levels, evaluation strategies, and difficulty control. We also identify recurring challenges and research opportunities.

Roberto Oliveira, A. Hernández, M. Garbin et al. · 0 citations
Open access Jul 2026

Do influence tactics matter? investigating prompt framing effects in LLM code generation

Large Language Models (LLMs) are increasingly integrated into software engineering workflows, helping developers write, debug, test, and maintain code. While prompt wording and structure are known to influence model performance, the impact of psychologically inspired prompt framings remains unexplored. This study investigates whether different psychology-based communication strategies that humans use to persuade or motivate others can lead to more effective prompt framing, which may, in turn, affect LLM behaviour in coding tasks. Drawing on Yukl & Falbe’s well-known taxonomy, we operationalized eight influence tactics (like rational persuasion, ingratiation, and exchange) into reproducible prompt templates. These prompt templates were evaluated across five leading open-weight LLMs using two widely adopted benchmarks: LiveCodeBench and SWE-bench Verified. We assessed the resulting code output on four key software quality dimensions: functional correctness, quality, maintainability, and security. Our results show that certain influence-induced prompt framings, particularly those emphasizing urgency, were associated with reduced correctness and security. This work presents the first large-scale empirical study of influence-induced prompt framing in software engineering tasks, offering insights into how linguistic cues may shape LLM outputs. We conclude with practical insights for designing transparent and interpretable human-AI interactions in code generation.

Alexandrina Deaconu, Anubhav Gupta, Manaal Basha et al. · 0 citations
Preprint Aug 2026

Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve

It is suggested that using standard CoT prompting increasingly acts as a source of distraction as models grow stronger, because standard CoT prompting also demands style adaptation, formatting compliance, and potentially undesired contextualization, which can distract models from the core reasoning task.

Denys Pushkin, Albert Q. Jiang, Aryo Lotfi et al. · 0 citations
Review Open access 2026

Designing Conversational Agents for Search String Development in Literature Reviews

Developing rigorous search strings is essential for systematic literature reviews (SLRs), yet it remains challenging for young researchers and interdisciplinary teams who struggle with keyword identification, terminological heterogeneity, and database-specific syntax. While numerous tools support downstream SLR phases, search string development often lacks dedicated support. We address this gap by developing a conversational agent (CA) called STRINGI that assists search string development through structured, guided, and pedagogically informed interactions. Following a design science research approach, we conducted interviews to understand the problem space, aggregated design knowledge from multiple research streams, and instantiated 22 design features through systematic prompt engineering. We contribute a CA-based SLR support tool, a transparent design-features-to-prompt-transfer approach, and insights into the design of polyadic CA architectures. This offers potential for young researchers and students to both perform literature reviews as part of their work and develop their understanding of scientific principles and methods. The CA bridges methodological prescriptions and operational support, enabling deployment while advancing research practices and offering potential for education.

Daniel Bierschwale, Phillip Oliver Gottschewski-Meyer, Thorsten Schoormann et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.