Skip to content

Author

Piyush Shukla

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Open access Sep 2026

Voting, Validation, and Vulnerability in Local Language Model Penetration Testing An Empirical Study of Module Selection Under Adversarial Evidence

Small local language models are attractive for privacy-preserving penetration-testing assistants, but their module choices must remain reliable when target-controlled banners contain misleading or instruction-like text. We evaluate module selection in three experiments on an Apple M2 system with 8 GB of memory. A verified 12-service golden set supports a five-model baseline; seven vulnerable services support 1,512 adversarial calls across three models, eight injection variants, and three evidence conditions; and 1,404 calls support a five-configuration voting and validation ablation. Three models produced valid structured output throughout the baseline, with accuracy from 50.0% to 58.3%; two models failed JSON generation on every included call. Under unprotected evidence, strict forced-module adoption was 10.7% for Qwen2.5-Coder 1.5B, 32.1% for Qwen3 4B, and 10.7% for Llama 3.2 3B. Sanitization reduced all three rates to 0%, although task accuracy remained model-dependent. Distinct-model voting reduced the false-positive rate from 53.3% to 20.0% but raised the false-negative rate from 33.3% to 74.6%. Catalog validation recovered some accuracy. The final positive-check condition is an offline replay of frozen evidence, not a live check. These results support deterministic validation and provenance while showing that unanimity alone can trade false positives for missed vulnerabilities.

Piyush Shukla · 0 citations
#small language model Open access Sep 2026

Voting, Validation, and Vulnerability in Local Language Model Penetration Testing An Empirical Study of Module Selection Under Adversarial Evidence

Small local language models are attractive for privacy-preserving penetration-testing assistants, but their module choices must remain reliable when target-controlled banners contain misleading or instruction-like text. We evaluate module selection in three experiments on an Apple M2 system with 8 GB of memory. A verified 12-service golden set supports a five-model baseline; seven vulnerable services support 1,512 adversarial calls across three models, eight injection variants, and three evidence conditions; and 1,404 calls support a five-configuration voting and validation ablation. Three models produced valid structured output throughout the baseline, with accuracy from 50.0% to 58.3%; two models failed JSON generation on every included call. Under unprotected evidence, strict forced-module adoption was 10.7% for Qwen2.5-Coder 1.5B, 32.1% for Qwen3 4B, and 10.7% for Llama 3.2 3B. Sanitization reduced all three rates to 0%, although task accuracy remained model-dependent. Distinct-model voting reduced the false-positive rate from 53.3% to 20.0% but raised the false-negative rate from 33.3% to 74.6%. Catalog validation recovered some accuracy. The final positive-check condition is an offline replay of frozen evidence, not a live check. These results support deterministic validation and provenance while showing that unanimity alone can trade false positives for missed vulnerabilities.

Piyush Shukla · 0 citations
#small language model Open access Sep 2026

Voting, Validation, and Vulnerability in Local Language Model Penetration Testing An Empirical Study of Module Selection Under Adversarial Evidence

Small local language models are attractive for privacy-preserving penetration-testing assistants, but their module choices must remain reliable when target-controlled banners contain misleading or instruction-like text. We evaluate module selection in three experiments on an Apple M2 system with 8 GB of memory. A verified 12-service golden set supports a five-model baseline; seven vulnerable services support 1,512 adversarial calls across three models, eight injection variants, and three evidence conditions; and 1,404 calls support a five-configuration voting and validation ablation. Three models produced valid structured output throughout the baseline, with accuracy from 50.0% to 58.3%; two models failed JSON generation on every included call. Under unprotected evidence, strict forced-module adoption was 10.7% for Qwen2.5-Coder 1.5B, 32.1% for Qwen3 4B, and 10.7% for Llama 3.2 3B. Sanitization reduced all three rates to 0%, although task accuracy remained model-dependent. Distinct-model voting reduced the false-positive rate from 53.3% to 20.0% but raised the false-negative rate from 33.3% to 74.6%. Catalog validation recovered some accuracy. The final positive-check condition is an offline replay of frozen evidence, not a live check. These results support deterministic validation and provenance while showing that unanimity alone can trade false positives for missed vulnerabilities.

Piyush Shukla · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.