Skip to content

Decision Hijacking: Prompt Injection Attacks on Jev's Typed Probabilistic Decisions

Sep 2026 · 5 citations · 11 references
Computer Science

TL;DR

It is shown that schema-defined outputs change but do not eliminate prompt-injection risk, highlighting the need to evaluate how untrusted content influences choices within the allowed action set.

Abstract

Most studies of prompt injection focus on generative agents, leaving their effects on models with schema-defined outputs unclear. We examine these effects in Jev, a non-generative decision model, using 510 reconstructed InjecAgent cases. Malicious content shifts action probabilities but rarely causes Jev to select the attacker's target. Override markers reduce this influence, while claims of contextual relatedness have small effects. Adaptive attacks using score feedback double the mean highest attacker-target probability found during optimization, while success on fresh validation calls rises from 1.8% to 3.5%. Exploratory analysis links these successes to small initial decision margins or greater attacker control over the observation. Together, these findings show that schema-defined outputs change but do not eliminate prompt-injection risk, highlighting the need to evaluate how untrusted content influences choices within the allowed action set.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Decoding Guardrails: XAI-Guided Perturbation Analysis of Prompt Injection Detection

Large language models (LLMs) are increasingly deployed in production systems, raising concerns about their exposure to adversarial manipulation through prompt injection and jailbreak attacks. Classifier-based guardrails, such as Prompt Guard 2, are widely used as a first line of defense against such attacks, but their...

Fernando Outeda, Gustavo Betarte, J. Campo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Rethinking Indirect Prompt Injection as a Test-Time Search Problem

This work identifies the attacker's adaptive search over the system attack surfaces as an important and underexplored security risk for tool-using agents and suggests that agentic security evaluations should characterize both the attacker's search procedure and compute budget.

D. M. Nguyen, Joon Sik Kim, Blazej Manczak et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents

Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skill...

T. Kaiser, Aritra Dhar · 0 citations
#machine learning Review Sep 2026

MOLE: Detecting Insider Threats in AI Agents

Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introd...

Aashiq Muhamed, Virginia Smith · 0 citations
Preprint Aug 2026

When Detection Does Not Guarantee Resistance: Reasoning and Poisoned Context in RAG

Retrieval-Augmented Generation (RAG) exposes large language models to knowledge-poisoning attacks, where misinformation injected into retrieved documents can influence model outputs. Prior work has shown that models may detect contradictory evidence yet still allow it to influence their responses, revealing a gap betwe...

Mehrdad Ghassabi, Audrina Ebrahimi, Sadra Hakim et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.