Skip to content
Open access

Comparative Evaluation of Prompt Engineering Configurations and Open-Source Large Language Models for Phishing Email Detection

2026 · International journal of research and innovation in social science · 0 citations

TL;DR

Overall, the results show that prompt configuration and model choice jointly influence phishing detection performance, false-alarm rates, output reliability, and local processing efficiency.

Abstract

Large Language Models (LLMs) offer a flexible approach to phishing-email classification, but their performance can vary with both prompt design and model selection. This study evaluates these factors through a two-stage experimental design using a balanced dataset of 1,000 emails comprising 500 phishing and 500 legitimate messages. In Stage 1, five complete prompt-engineering configurations —Zero-Shot, Few-Shot Demonstration, Expert Role-Based, Structured Criteria-Guided, and CoT-Style Structured Reasoning Prompting— were compared using Mistral. The best-performing configuration was then held constant in Stage 2 to compare five locally deployed open-source LLMs: Mistral, Phi-3, Qwen2.5, Llama 3.1, and Gemma 2. Performance was assessed using accuracy, precision, recall, F1-score, valid-output coverage, and processing time. Structured Criteria-Guided Prompting achieved the strongest prompt-level performance, with 86.33% accuracy and an F1-score of 85.43%. In the model comparison, Gemma 2 achieved the highest accuracy (87.00%), recall (90.00%), and F1-score (87.38%), while maintaining 100% valid-output coverage. Qwen2.5 achieved the highest precision (95.73%) and the shortest recorded processing time of 1 h 7 min 46 s. Valid-output coverage also varied across prompt configurations, with Expert Role-Based Prompting producing the lowest coverage at 88.2%. Overall, the results show that prompt configuration and model choice jointly influence phishing detection performance, false-alarm rates, output reliability, and local processing efficiency.

Read PDF

Similar papers

Preprint Aug 2026

Aging of Prompt Engineering Techniques Across LLM Versions

It is shown that prompt engineering "ages" in a model-family-specific way: Newer GPT models exhibit diminishing or even negative marginal gains from structured prompting, suggesting that instruction-following and reasoning scaffolds are increasingly internalized, whereas Qwen models continue to benefit substantially fr...

A. Rudyk, Julian Oertel, Regina Hebig · 0 citations
Review Open access Sep 2026

Evaluating large language models as grant reviewers: a comparative study of prompt engineering strategies

Grant application review is resource-intensive and subject to inter-rater variability. Large language models (LLMs) may augment this process, but their reliability in grant evaluation remains unexplored. This exploratory pilot study compared LLM-generated grant reviews to human expert reviews across three prompt engi...

Hants Williams, Jack Evan Lamberg, Eric M. Lamberg · 0 citations
Conference Aug 2026

When Iterative Prompting Fails: An Empirical Study of Unit Test Generation with Open-Source LLMs

Large language models (LLMs) have shown promise in automated unit test generation, yet the effectiveness of prompt engineering for small, locally-deployed open-source models remains poorly understood. Following growing interest in local LLM deployment to mitigate data exposure risks, this paper presents a controlled em...

M. Tran, Khang Mai · 0 citations
#small language model Preprint Sep 2026

Configuration, Not Conscience: A Large-Scale Empirical Study of LLM System Prompts

Analysis of leaked system prompts from 62 vendors across four community collections supports treating leaked prompts as operational specifications, closer to configuration files than value statements, and treats reuse and prompt rot as engineering and supply-chain concerns.

C. Patsakis, Vasilios Argyropoulos, Efthymios Alepis · 0 citations
Open access Sep 2026

Evaluating Prompt Engineering Techniques for LLaMA-3: A Study of Zero-Shot, Few-Shot, and Chain-of-Thought Prompts Across Reasoning and Classification Tasks

Prompt engineering has emerged as a practical and resource-efficient alternative to fine-tuning large language models (LLMs), particularly as these methods have a lower computation cost than fine-tuning. In this paper, three widely adopted prompting techniques—Zero-Shot, Few-Shot, and Chain-of-Thought (CoT)—were assess...

Darren Astle Travasso, Aboozar Taherkhani · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.