Skip to content
Review

SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets

Jul 2026 · arXiv.org · Vol abs/2607.08681 · 0 citations · 35 references
Computer Science Economics

TL;DR

SolarChain-Eval is a physics-constrained benchmark for evaluating trustworthy economic agents and formulates market governance as a Gymnasium-compatible Markov Decision Process, where agents make hourly decisions.

Abstract

As agentic AI systems are increasingly applied to cyber-physical environments, their evaluation requires assessment of both task performance and trustworthiness. In decentralized energy markets, autonomous agents may improve market utility, but may also exploit invalid physical data, create artificial liquidity, and produce unstable governance decisions. Therefore, we propose SolarChain-Eval, a physics-constrained benchmark for evaluating trustworthy economic agents. It formulates market governance as a Gymnasium-compatible Markov Decision Process, where agents make hourly decisions. SolarChain-Eval evaluates each policy across multiple dimensions, including market utility, physical safety, slippage, action smoothness, spatial fairness, and auditability. To support agentic evaluation, SolarChain-Eval incorporates an LLM-based Planner/Auditor layer. The Planner defines episode-level action bounds and audit rules, while the Auditor reviews and revises high-risk actions. All interventions are recorded through structured logs, including trigger signals, proposed actions, revised actions, and audit rationales. Experiments with static, random, myopic, RL, and RL+LLM policies reveal a clear utility-safety trade-off. RL agents improve market utility but can still produce unsafe behavior. When the physics penalty is removed, reward-maximizing agents exploit invalid generation and increase artificial liquidity. The LLM Planner/Auditor improves auditability and mitigates selected risks, but it cannot fully compensate for a misspecified reward function. These results indicate that trustworthy agentic AI evaluation requires both physical constraints and transparent intervention traces. We release data and code as open access on GitHub for replicability.

View source

Similar papers

Preprint Aug 2026

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

Business Arena, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon, is introduced, taking a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.

Yijun Pan, Yu-Kun Lian, Kun-Yu Shi et al. · 1 citation
Conference Aug 2026

COALITION-VAST: Auditable Multi-Agent Alignment Under Byzantine Governance

Scaling aligned AI from single-agent systems to multi-agent ecosystems introduces collective failures that do not arise in isolation: coalition deviation, governance capture, and rushed rule changes. Prior work in VAST and VAST-Blockchain addresses single-agent compliance and deployment integrity, but not strategic coo...

Soraya Partow, Satyaki Nan · 0 citations
Conference Jul 2026

The Agentic AI Framework for Optimizing Yield Aggregators in Decentralized Finance

As decentralized finance (DeFi) ecosystems continue to expand, yield aggregators play an important role in automating capital allocation across lending and liquidity protocols. However, most existing aggregators still rely on static strategies and governance-driven update cycles, limiting their ability to respond to ra...

Alvan Nauval, A. Alamsyah · 0 citations
Preprint Aug 2026

A-CPES: A Reference Framework for Agentic AI in Cyber-Physical Energy Systems

It is argued the loop is indivisible, tune where and how tightly it may close, state eight structural failure modes as falsifiable predictions, and specify six governance modules that rebuild the authorization frame until it covers the loop, before the loop starts turning.

Xiaoyu Zhang, Qiu-Ye Sun, Jiachen Xu et al. · 0 citations
Preprint Aug 2026

Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems

A controlled, physics-grounded benchmark built around planning-induced control trajectories: the ordered planning operations and directives through which an execution architecture acts on other agents and the physical process is introduced.

J. de Curtò, I. de Zarzà · 1 citation
Sep 2026

A unified architecture for sanctions intelligence in digital asset markets

Sanctions compliance in digital asset markets raises challenges that differ materially from those encountered in traditional financial systems. Public blockchains introduce pseudonymity, indirect exposure, timing effects, and a rapidly evolving landscape of cross-chain movement and decentralised finance activity that c...

Aniket Mandavkar, Tony Gagliardi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.