What the results imply for how such orchestrators should be evaluated are discussed, and the study uses a single benchmark, the fixes were tuned on the same challenges they were evaluated on, and the client effect is demonstrated for one model only, so its generality to other models remains a hypothesis.
Abstract
Large language model agents driving security tool suites over the Model Context Protocol are increasingly common. Yet the factors that bound their capability remain poorly characterized: how much depends on the model versus the client that drives it, whether constraining the agent to the orchestrator's own tools helps, and where capability is limited by reasoning rather than by missing tools. Using HexStrikeAI, an open-source orchestrator that exposes 150+ tools, as a testbed, we follow a methodology that evaluates the system, diagnoses its failures, and applies targeted improvements. We run 86 picoCTF challenges across seven categories and three difficulty tiers, under three tool-access regimes and three model/client configurations (774 trials). We then apply corrections to existing tools, agent-behavior changes, and eleven new capability tools, and re-run the previously-unsuccessful trials. The diagnosis isolates the driving client as a first-order factor for a fixed model (a 2.1 * gap between two DeepSeek clients) and a monotonic difficulty gradient, with the largest gains in the mid tier. The overall solve rate rises from 55.4% to 72.0%, and every configuration improves significantly (paired McNemar p<0.001, non-overlapping 95% confidence intervals). The residual failures are reasoning- or environment-bound rather than missing-tool. A 60-run stability sub-study finds single-run verdicts reproducible (17/20 unanimous). We discuss what the results imply for how such orchestrators should be evaluated, and we are explicit about the limits: the study uses a single benchmark, the fixes were tuned on the same challenges they were evaluated on, and the client effect is demonstrated for one model only, so its generality to other models remains a hypothesis.
A four-dimensional Integration Friction Index is introduced that separates one-time engineering cost from recurring organisational, legal, and maintenance cost and shows why scope and budget enforcement cannot be delegated to system prompts.
Israt Moyeen Noumi, Tarannum Ahmed Nowshin, Md. Mehedi Hasan Bhuiyan Nipu et al.· 0 citations
This paper presents ToolGuardian, a policy-driven framework for securing agent-tool interactions through pre-admission vetting and task-aware runtime authorization, and compares ASP against heuristic and LLM-based policy realizations using identical inputs and output contracts.
Raising effort did change behaviour, but only in inspection: rule-probe rates rose in all conditions, but only in inspection: rule-probe rates rose in all conditions, a pattern inconsistent with the hypothesis of targeted search.
This survey treats isolation as a first-class principle for LLM-agent system safety, and organizes the literature with a boundary-centric taxonomy of five boundaries: user-agent, agent-tool, agent-execution, agent-agent, and system-environment.
Context: The growing complexity of cloud microservices imposes significant challenges for Site Reliability Engineering (SRE), contributing to delayed incident resolution and increased operational effort. Objective: This study evaluated the effectiveness of autonomous agents based on Large Language Models (LLMs), orchestrated via the Model Context Protocol (MCP), for root cause analysis in a cloud-native setting. Method: We conducted a controlled Randomized Complete Block Design (RCBD) experiment in Kubernetes with automated fault injection, covering three distinct failure scenarios and multiple LLM configurations across 360 executions. Results: A high-performing configuration (Gemini 2.5 Flash at low temperature) achieved a 71.1% root-cause identification success rate, substantially above a random-chance baseline (≈ 0.91%). Smaller models exhibited higher token and step volatility (CV = 2.17) and more repeated tool-call cycles, challenging the assumption that lower-parameter models are inherently more cost-effective for SRE workflows. Conclusion: The results provide empirical evidence that MCP-orchestrated LLM agents can support root cause analysis in cloud-native environments and offer practical guidance for model selection in AIOps/SRE workflows.
Reinan Gabriel dos Santos Souza, Methanias Colaço· Anais do LIII Seminário Inte...· 0 citations
This work argues that agentic risk is progressive: it can enter at four loci of the agent control loop--skill admission, invocation-time intent, execution-time effect, and post-action consequence--while a denied dangerous objective can reappear across surface forms, tools, or turns.
Kai Wang, Zeming Wei, Biaojie Zeng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.