Skip to content
Open access

Prompt Escalation for Lightweight Large Language Models: An Empirical Evaluation of Cost–Performance Trade-Offs

Sep 2026 · Applied Sciences · 0 citations · 18 references

TL;DR

Results support ZS as a low-overhead reference within the four evaluated benchmarks and the stated model, quantization, prompt, and decoding settings, with model–task exceptions.

Abstract

Prompt escalation can increase the resource requirements of lightweight large language models (LLMs) without improving predictive performance. We evaluated four instruction-tuned 2–4B models on Grade School Math 8K (GSM8K), CommonsenseQA (CSQA), Recognizing Textual Entailment (RTE), and the binary Stanford Sentiment Treebank (SST-2). Zero-shot (ZS), few-shot (FS), chain-of-thought (CoT), and few-shot CoT (FS+CoT) were compared across 64 conditions using predictive performance, tokens, latency, and stochastic response consistency. Demonstrations came from training splits; FS+CoT used worked rationales with automatic screening and a partial manual audit. Latency was measured separately with synchronization, warm-up exclusion, and counterbalanced prompt order. After Holm adjustment, 12.5% of predictive-performance contrasts were significant, and ZS was significantly outperformed in 1 of 48 contrasts, compared with significant differences in 100% of total-token and 75.0% of latency contrasts. Strategy rankings varied by model and task. A single auxiliary 7B model showed no significant accuracy gain over ZS but did not establish a general scale effect. Scenario-weight sensitivity frequently favored ZS, with model–task exceptions. These results support ZS as a low-overhead reference within the four evaluated benchmarks and the stated model, quantization, prompt, and decoding settings. The rankings and guidelines have not been validated for summarization, code generation, or multi-turn dialogue.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation

This work investigates LLM-based evaluators of natural language generation quality mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and expli...

Himil Vasava, Ming-Zhou Jiang · 0 citations
Open access Sep 2026

Evaluating Prompt Engineering Techniques for LLaMA-3: A Study of Zero-Shot, Few-Shot, and Chain-of-Thought Prompts Across Reasoning and Classification Tasks

Prompt engineering has emerged as a practical and resource-efficient alternative to fine-tuning large language models (LLMs), particularly as these methods have a lower computation cost than fine-tuning. In this paper, three widely adopted prompting techniques—Zero-Shot, Few-Shot, and Chain-of-Thought (CoT)—were assess...

Darren Astle Travasso, Aboozar Taherkhani · 0 citations
Open access Aug 2026

Empowering Small Language Models With Chain of Thought and Parameter Efficient Fine Tuning for Efficient Deep Reasoning

: Small Language Models (SLMs) run faster and fit on modest hardware, yet solving multi-step logic problems has traditionally been difficult for them. This work investigates a systematic, multi-model framework that combines Chain-of-Thought (CoT) prompting with Low-Rank Adaptation (LoRA) parameter-efficient fine-tuning...

Aryan Raina, Shiwani Gupta, Jagruti Jadhav et al. · 0 citations
Open access 2026

Comparative Evaluation of Prompt Engineering Configurations and Open-Source Large Language Models for Phishing Email Detection

Overall, the results show that prompt configuration and model choice jointly influence phishing detection performance, false-alarm rates, output reliability, and local processing efficiency.

Naavin A/L Rajeswaren, Agus Hartoyo, Farhad Nadi · 0 citations
#natural language process... Preprint Sep 2026

Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation

Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spannin...

Mario Sanz-Guerrero, M. Bui, Manuel Mager et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.