Small local language models are attractive for privacy-preserving penetration-testing assistants, but their module choices must remain reliable when target-controlled banners contain misleading or instruction-like text. We evaluate module selection in three experiments on an Apple M2 system with 8 GB of memory. A verified 12-service golden set supports a five-model baseline; seven vulnerable services support 1,512 adversarial calls across three models, eight injection variants, and three evidence conditions; and 1,404 calls support a five-configuration voting and validation ablation. Three models produced valid structured output throughout the baseline, with accuracy from 50.0% to 58.3%; two models failed JSON generation on every included call. Under unprotected evidence, strict forced-module adoption was 10.7% for Qwen2.5-Coder 1.5B, 32.1% for Qwen3 4B, and 10.7% for Llama 3.2 3B. Sanitization reduced all three rates to 0%, although task accuracy remained model-dependent. Distinct-model voting reduced the false-positive rate from 53.3% to 20.0% but raised the false-negative rate from 33.3% to 74.6%. Catalog validation recovered some accuracy. The final positive-check condition is an offline replay of frozen evidence, not a live check. These results support deterministic validation and provenance while showing that unanimity alone can trade false positives for missed vulnerabilities.
Piyush Shukla· Zenodo (CERN European Organi...· 0 citations
Small local language models are attractive for privacy-preserving penetration-testing assistants, but their module choices must remain reliable when target-controlled banners contain misleading or instruction-like text. We evaluate module selection in three experiments on an Apple M2 system with 8 GB of memory. A verified 12-service golden set supports a five-model baseline; seven vulnerable services support 1,512 adversarial calls across three models, eight injection variants, and three evidence conditions; and 1,404 calls support a five-configuration voting and validation ablation. Three models produced valid structured output throughout the baseline, with accuracy from 50.0% to 58.3%; two models failed JSON generation on every included call. Under unprotected evidence, strict forced-module adoption was 10.7% for Qwen2.5-Coder 1.5B, 32.1% for Qwen3 4B, and 10.7% for Llama 3.2 3B. Sanitization reduced all three rates to 0%, although task accuracy remained model-dependent. Distinct-model voting reduced the false-positive rate from 53.3% to 20.0% but raised the false-negative rate from 33.3% to 74.6%. Catalog validation recovered some accuracy. The final positive-check condition is an offline replay of frozen evidence, not a live check. These results support deterministic validation and provenance while showing that unanimity alone can trade false positives for missed vulnerabilities.
Piyush Shukla· Zenodo (CERN European Organi...· 0 citations
Small local language models are attractive for privacy-preserving penetration-testing assistants, but their module choices must remain reliable when target-controlled banners contain misleading or instruction-like text. We evaluate module selection in three experiments on an Apple M2 system with 8 GB of memory. A verified 12-service golden set supports a five-model baseline; seven vulnerable services support 1,512 adversarial calls across three models, eight injection variants, and three evidence conditions; and 1,404 calls support a five-configuration voting and validation ablation. Three models produced valid structured output throughout the baseline, with accuracy from 50.0% to 58.3%; two models failed JSON generation on every included call. Under unprotected evidence, strict forced-module adoption was 10.7% for Qwen2.5-Coder 1.5B, 32.1% for Qwen3 4B, and 10.7% for Llama 3.2 3B. Sanitization reduced all three rates to 0%, although task accuracy remained model-dependent. Distinct-model voting reduced the false-positive rate from 53.3% to 20.0% but raised the false-negative rate from 33.3% to 74.6%. Catalog validation recovered some accuracy. The final positive-check condition is an offline replay of frozen evidence, not a live check. These results support deterministic validation and provenance while showing that unanimity alone can trade false positives for missed vulnerabilities.
Piyush Shukla· Zenodo (CERN European Organi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.