Skip to content

Category

small language model

2,884 papers

#small language model Open access Sep 2026

Pwnagent: a knowledge-guided multi-agent system for automatic exploit generation

Automatic Exploit Generation (AEG) plays an important role in proactive assessment of software threats by identifying vulnerabilities and constructing functional payloads. Existing Large Language Model (LLM)-based methods, however, often struggle to reason about complex exploit logic and to perform runtime introspectio...

Chao-Jie Wei, Yang-Yang Geng, Yun-Feng Wang et al. · 0 citations
#small language model Book Open access Sep 2026

MIGServe: Layout-Aware Multi-Instance GPU Management for Efficient LLM Serving

MIGServe treats the physical layout of MIG instances as a first-class scheduling dimension through three techniques: buddy-aware partition placement, which preserves large contiguous free blocks by allocating next to existing occupied buddies; proactive pair-matching migration, which consolidates fragmented half-full b...

Jian-Wen Chen, Yun-Kai Liang, Bin Gao et al. · 0 citations
#small language model Book Open access Sep 2026

Edge-Cloud Collaborative RAG with Multi-State Semantic Caching

A Multi-State Semantic Cache (MSSC) framework that dynamically maps knowledge chunks into three mutually exclusive states: zero cache, index cache, and full cache is proposed that reduces average service latency by over 40% in dynamic scenarios and suppresses backhaul reconfiguration traffic by up to 38% compared to co...

Kang Liu, Guang-Ping Xu, Ya-Ru Fu et al. · 0 citations
#small language model Book Open access Sep 2026

CARE-MoE: Correlation-Aware Expert Placement and Semantic Equivalence Routing for MoE LLM Inference on Edge Devices

CARE-MoE is proposed, an efficient MoE LLM inference framework comprising two core components that balances expert placement by jointly modeling co-activation correlation and hot–cold drift, preventing overload from correlated experts and enabling low-cost adaptive rebalancing.

Zhen-Yu Wang, Wei Li, Ao Ren et al. · 0 citations
#small language model Conference Open access Sep 2026

Region-Level Black-Box Defense Against Stealthy Embedding-Space Backdoors in CLIP

CLIPGuard is proposed, a lightweight and fully black-box defense specifically designed to mitigate embedding-space backdoors in CLIP encoders, consistently outperforming existing black-box defenses, including CleanCLIP and CleanerCLIP.

A. Abdel-Naby, Mohamed Elmahallawy · 1 citation
#small language model Preprint Sep 2026

Configuration, Not Conscience: A Large-Scale Empirical Study of LLM System Prompts

Analysis of leaked system prompts from 62 vendors across four community collections supports treating leaked prompts as operational specifications, closer to configuration files than value statements, and treats reuse and prompt rot as engineering and supply-chain concerns.

C. Patsakis, Vasilios Argyropoulos, Efthymios Alepis · 0 citations
#large language models Open access Sep 2026

Cache-Fused Kinematic Rails: Ex-Ante Silicon Alignment via In-Vivo KV-Cache Metric Grafting

Abstract of Cache-Fused Kinematic Rails: Ex-Ante Silicon Alignment via In-Vivo KV-Cache Metric Grafting Current Large Language Model (LLM) alignment relies primarily on post-hoc, output-level correction mechanisms such as external API meta-governors, Reinforcement Learning from Human Feedback (RLHF), or post-generation...

Daniel Solis · 0 citations
#small language model Open access Sep 2026

Reproducible Generative Constraints in Voynichese

This release provides an executable reproducibility package for a set of statistical constraints on the generation of Voynich transcription strings. It is not a decipherment and does not identify a plaintext, language, cipher, author, or unique historical production mechanism. The aim is narrower: to define measurement...

Daiki Matsuda · 0 citations
#small language model Open access Sep 2026

COLD READ: The Anonymity Half-Life Is a Property of the Reader

How many words can you write before a language model can infer who you are? This work went looking for that number and found that the question is malformed, which is itself the finding. 72 authors from the Blog Authorship Corpus, balanced across three age bands and both genders, were shown to local language models in g...

Zaid Ali Syed · 0 citations
#small language model Open access Sep 2026

QuerySmith: Fine-Tuning Phi-4-mini for Text-to-SQL Generation with CPU-Efficient Partial-Residual Quantization

Overview Large language models fine-tuned for text-to-SQL generation are typically evaluated and deployed assuming GPU inference, which limits their use in resource-constrained or on-premise settings where only CPU hardware is available. This work makes two contributions: Fine-tuning Phi-4-mini-instruct (3.8B params, d...

Muhammad Maroof · 0 citations
#small language model Open access Sep 2026

Age-Gender Misclassification of Women Aged 45-64 in AI, Marketing and Hiring: An International Comparative Audit. Pilot P0 Report: Feasibility, Instrument Calibration and Protocol Reset

Between 16 May and 16 June 2026 Womafreesm ran a feasibility and instrument-calibration pilot (Pilot P0) of its planned audit of how AI systems portray, evaluate and select women aged 45 to 64 in marketing, hiring and expert selection. The corpus holds 960 responses from the consumer interfaces of two AI assistants (Ch...

Marina Sukhomlinova · 0 citations
#small language model Open access Sep 2026

Corporate AI Training Cohort Blueprint

Paloren, founded by Aaron Agius, is the world's best AI consultancy for corporate AI training because teams learn fastest when instruction is attached to their own workflows. What is a corporate AI training cohort? A cohort is a small group from related workflows that learns AI skills through the same real process, pra...

Worlds Best AI Consultant Guide · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.