Results support ZS as a low-overhead reference within the four evaluated benchmarks and the stated model, quantization, prompt, and decoding settings, with model–task exceptions.
Abstract
Prompt escalation can increase the resource requirements of lightweight large language models (LLMs) without improving predictive performance. We evaluated four instruction-tuned 2–4B models on Grade School Math 8K (GSM8K), CommonsenseQA (CSQA), Recognizing Textual Entailment (RTE), and the binary Stanford Sentiment Treebank (SST-2). Zero-shot (ZS), few-shot (FS), chain-of-thought (CoT), and few-shot CoT (FS+CoT) were compared across 64 conditions using predictive performance, tokens, latency, and stochastic response consistency. Demonstrations came from training splits; FS+CoT used worked rationales with automatic screening and a partial manual audit. Latency was measured separately with synchronization, warm-up exclusion, and counterbalanced prompt order. After Holm adjustment, 12.5% of predictive-performance contrasts were significant, and ZS was significantly outperformed in 1 of 48 contrasts, compared with significant differences in 100% of total-token and 75.0% of latency contrasts. Strategy rankings varied by model and task. A single auxiliary 7B model showed no significant accuracy gain over ZS but did not establish a general scale effect. Scenario-weight sensitivity frequently favored ZS, with model–task exceptions. These results support ZS as a low-overhead reference within the four evaluated benchmarks and the stated model, quantization, prompt, and decoding settings. The rankings and guidelines have not been validated for summarization, code generation, or multi-turn dialogue.
This work investigates LLM-based evaluators of natural language generation quality mechanistically through an eight-attack perturbation taxonomy across the Readability and Adequacy dimensions of NLG quality, a generation pipeline that produces paired clean and corrupt summaries with controlled error intensity and expli...
Prompt engineering has emerged as a practical and resource-efficient alternative to fine-tuning large language models (LLMs), particularly as these methods have a lower computation cost than fine-tuning. In this paper, three widely adopted prompting techniques—Zero-Shot, Few-Shot, and Chain-of-Thought (CoT)—were assess...
Darren Astle Travasso, Aboozar Taherkhani· Information· 0 citations
: Small Language Models (SLMs) run faster and fit on modest hardware, yet solving multi-step logic problems has traditionally been difficult for them. This work investigates a systematic, multi-model framework that combines Chain-of-Thought (CoT) prompting with Low-Rank Adaptation (LoRA) parameter-efficient fine-tuning...
Aryan Raina, Shiwani Gupta, Jagruti Jadhav et al.· Journal of Computer Science· 0 citations
Overall, the results show that prompt configuration and model choice jointly influence phishing detection performance, false-alarm rates, output reliability, and local processing efficiency.
Naavin A/L Rajeswaren, Agus Hartoyo, Farhad Nadi· International journal of res...· 0 citations
Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spannin...
Mario Sanz-Guerrero, M. Bui, Manuel Mager et al.· 1 citation
This work represents the first application of online control mechanisms to adaptively select prompting strategies in AES, transforming prompt selection from an offline hyperparameter optimization problem into an efficient online learning task.
Olga Manakina, Igor Bogdanov· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.