2026· International Journal of Scientific Research and Management· Vol 14, pp. 2771-2780· 0 citations
TL;DR
Intent-Based Security (IBS), a structured approach built on foundational ideas from access control, zero-trust models, and principal-agent dynamics, shows why trusting identities fails against invisible threats.
Abstract
Earlier work showed clean attacks succeed about 61.4% of the time when targeting advanced LLM agents in one-step interactions; performance climbs to 90.0% by the twentieth step if constraints weaken gradually. These findings share a core insight: current defenses rely too heavily on where inputs come from, not what they aim to do. Because trust travels with verified sources, malicious additions slip through unnoticed once embedded steering systems off course without raising alarms. From this stems the risk: legitimacy of origin masks harmful intent buried inside otherwise trusted material. This work shows why trusting identities fails against invisible threats. Even when a message comes from a verified origin, meets every security rule, arrives encrypted, it might still carry hidden directions meant to mislead automated systems. The flaw lies not in who sent it, but in what the message means. Security checks based on user roles, token updates, or digital signatures do nothing here these tools ignore meaning entirely. Trust built on identity does not guarantee safe outcomes. What matters shifts from source validation to purpose matching. That mismatch defines the problem: proof of authenticity rarely equals proof of honest intent. Addressing this issue begins with Intent-Based Security (IBS), a structured approach built on foundational ideas from access control, zero-trust models, and principal-agent dynamics. Rather than relying solely on who sends data, trust now depends on how closely each incoming input matches the operator's stated intent vector φ checked continuously. This system unfolds across three central elements. One key part the Intent Verification Pipeline (IVP) acts as a layered checkpoint, producing a dynamic trust score τ for every input. That value emerges by combining insights from the Semantic Validity Score (SVS), detailed earlier in Paper 1, along with adjustments based on the Intent Drift Rate (IDR) from Paper 2. Periodically, Intent Anchoring brings φ back into the agent’s context every k turns, mirroring the restorative mechanism outlined in Paper 2's Lyapunov stability approach. Following that, the system introduces a boundary called φ-Trust Threshold τ, adjusted using Bayesian decision principles, which separates actions labeled execute from those marked block. One test after another, covering various attack types and four leading LLM systems, shows IBS using complete IVP setup and intent anchoring (k = 5) brings average ASR down from 61.4% to 11.6%, cutting it by 81.1 percentage points, yet keeps incorrect flags low at just 4.1%. Earlier alerts also emerge: about 3.2 turns ahead of major shifts in behavior seen earlier in Paper 2.
Zero-day attacks exploit signature-based defenses since the key pieces of evidence used to identify attacks are not present during the training of the model. This paper proposes an Agentic Autonomous Zero-Trust (AAZT) framework that includes: Multimodal Open-set Detection, Evidence-grounded Multi-agent Investigation, C...
Indirect prompt injection attacks - malicious instructions embedded in content processed by large language models - remain a major obstacle to safely deploying tool-using agents. CaMeL [Debenedetti et al., 2025] mitigates this threat for an individual agent by separating trusted control flow from untrusted data and enf...
James Peters-Gill, Avi Semler, Henning Bartsch et al.· 0 citations
This paper evaluated Niyam-AI on 2,000 real-world agent scenarios from Agent-SafetyBench and compared it against three existing safety approaches: NeMo Guardrails, Meta's Llama Prompt Guard 2, and OpenAI's GPT-OSS-Safeguard.
Agentic browsers can execute security-sensitive actions under a user's authenticated session, making indirect prompt injection and deceptive confirmation interfaces a direct threat to action integrity. Existing human-in-the-loop (HITL) safeguards are insufficient when the approval prompt itself can be influenced by unt...
Hasnain Irshad, Anam Mughees, Neelam Mughees et al.· 2 citations
This work argues that agents need an operating-system substrate providing mandatory, non-bypassable services for identity, input mediation, memory governance, and execution control, and introduces AgentKernel, a trust-native agent operating system built around the premise that security must be a first-class design cons...
Zhen-Hua Zou, Sheng Guo, Qiu-Yang Zhan et al.· 0 citations
Agentic artificial intelligence expands the enterprise security boundary because autonomous agents can plan
tasks, retain memory, invoke tools, call APIs, and initiate business actions. Authentication at session start is therefore
insufficient when later actions may be influenced by untrusted content, poisoned memory,...
S. Suryawanshi· International Journal of Inn...· 0 citations
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.
What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.