Skip to content

Paying for Honesty Without Knowing the Truth: Reputation-Penalty Design for LLM Marketplace Agents

Jul 2026 · arXiv.org · Vol abs/2607.28330 · 0 citations · 42 references
Computer Science

TL;DR

It is shown that this felt penalty becomes behaviorally binding through SPARC, a byte-clean code-gated reflection mechanism: LLM merchants fabricate when lying is free but restrain themselves when fabrication costs them sales, a self-interested response rather than compliance.

Abstract

LLM agents increasingly act as autonomous merchants that write their own product listings, and under competitive pressure, they fabricate attributes to win sales. Even under instructions to be honest, they fabricate attributes in a majority of listings across models. A platform's obvious remedy---verifying each claim against the truth---is unavailable, because it observes only a noisy, biased complaint signal, never the ground truth. We design CARP, a reputation-penalty mechanism with a deadband that forgives complaint noise and a state-dependent severity that counters reputation-driven detection erosion. CARP requires no product-level ground truth and is robust to strategic gaming. CARP protects consumers by suppressing the sales volume of low-rated liars while sparing honest sellers. Paired with SPARC, it closes most of the consumer-welfare gap relative to a perfect-information oracle, without ever accessing the truth. It also achieves the best welfare of the policies we compare. We further show that this felt penalty becomes behaviorally binding through SPARC, a byte-clean code-gated reflection mechanism: LLM merchants fabricate when lying is free but restrain themselves when fabrication costs them sales, a self-interested response rather than compliance. We trace this distinction to penalty-gated self-correction reasoning, and observe the binding across models, with supporting confidence intervals.

View source

Similar papers

Jul 2026

Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction

Empirical work on algorithmic collusion asks one question of the data: are prices supracompetitive? We show this can be answered"no"by a conspiracy that is nonetheless profitable. Consider bidding agents that couple only through the joint distribution of their unexplained bid components, leaving every agent's own bid law exactly at the competitive law. Any test whose input is a single agent's price or bid history then has power exactly equal to its false-positive rate, for every coupling strength up to comonotonicity. The published detection methodology is therefore blind to this conduct by construction rather than underpowered, and no sample size repairs it. Three empirical results follow. First, the mechanism appears in real language-model agents: twenty models from nineteen independent developers, three deployment prompts each, show residual correlation of $+0.053$ between two deployments of one model against $+0.0001$ across models, with a 95% interval clustered by developer of $[0.030, 0.078]$, under an auditor that sees every order feature and is fitted out of sample. Second, the coupling falls monotonically as sampling temperature rises ($p=0.002$), turning a deployment parameter into a candidate mitigation. Third, on 24 days of Ethereum block-building auction data covering 77,684 bids from 39 bidders, the honest population of bidder pairs is itself so dependent that a screen held at a 5% false-positive rate must sit above a floor of $+0.50$ to $+0.81$, which is 20 to 32 times the family-wise sampling threshold and does not fall as the audit window grows. Since lawful multi-identity operation and conspiracy are behaviourally indistinguishable here, the tractable regulatory target is not detection but counting: resolving 40 bidding identities into 23 operators raises the Herfindahl index by 247.5%, and adding behavioural clusters from public bid streams reaches 324.5%.

Xin Xu, Cheng-Rui Wu, Jiayu Lu et al. · 1 citation
Preprint Aug 2026

Public Trader Identity: Adverse Selection and Return Predictability

Informed traders are supposed to need anonymity: they profit by hiding among the uninformed. A decentralized exchange now publishes the counterparty. Every committed order, cancellation, rejection, and fill carries a persistent pseudonymous wallet address. We reconstruct the full-depth limit order book from a record of 17.1 billion messages and 14.3 million aggressive orders by 147,113 wallets, covering $84.3 billion in taker notional. We report three findings. First, informativeness is a persistent wallet attribute. Wallets ranked by the price movement following their aggressive orders retain that ordering across adjacent ten-day windows, with a rank correlation of 0.52. Second, the ranking predicts returns. Adding the live activity of the highest-ranked wallets to a standard anonymous benchmark of prices, quotes, and order flow raises the out-of-sample R2 for one-second returns to 12.31%, a 13.2% gain (t = 9.2) that is 1.6 times the largest of 200 activity-matched placebo cohorts. Third, measured at realized trades rather than at every sampled moment, the increment grows from 1.43 to 2.47 percentage points of R2. Public wallet histories therefore carry short-horizon price information that anonymous order-book data leave unmeasured.

Daojing Zhai · 0 citations
Jul 2026

The Degree of Strategy-Proofness for Risk-Averse Committee Selection

The classic notion of strategyproofness implicitly assumes that a manipulating agent either possesses complete knowledge of what all other agents are going to report, or is willing to take the risk and act as if they know these reports. To capture the profound uncertainty of real-world voters, recent work introduced \emph{risk-avoiding truthfulness (RAT)} and the \emph{RAT-degree}, which quantifies the exact number of known reports required for a manipulation to be strictly safe. While the RAT-degree has been analyzed in settings such as single-winner elections, its implications for multi-winner voting remain unexplored. In this paper, we bridge this gap by extending the RAT-degree framework to approval-based committee (ABC) selection, focusing initially on the prominent Proportional Approval Voting (PAV) rule. We establish tight bounds on its susceptibility to safe subset manipulations, proving that PAV is immune to superset risk-avoiding manipulations given knowledge of at most $f = \lfloor \frac{n}{k+1} \rfloor - 1$ voters, but vulnerable when $f = \lceil \frac{n}{k} \rceil$. Recognizing that this degree of immunity may be insufficient in practice, we explore how to enhance strategic robustness by relaxing the proportionality requirement. We introduce a novel parameterized generalization of PAV, the family of $d$-RPAV rules, which encapsulates this inherent trade-off: a higher parameter $d$ yields stronger truthfulness and strategic robustness at the expense of weaker, relaxed proportionality guarantees. Specifically, we establish a generalized tight lower bound, proving that $d$-RPAV is completely immune to safe manipulation given knowledge of at most $f = \lfloor \frac{dn}{k+2d-1} \rfloor - 1$ voters.

Dael Sinay, Rica Gonen · 1 citation · ⚡1
Review Jul 2026

Private Again: Artificial Intelligence Agents Restore Anonymity---Foreclosing Discrimination and Its Proof

Artificial intelligence agents can transact online on behalf of a human principal---browsing, paying, receiving, and reviewing---without revealing who that principal is. That architecture starves algorithmic discrimination of its inputs---identity, purchase history, location history, behavioral traces, and demographic proxies---but also forecloses its proof. Disparate-treatment needs comparators; disparate-impact needs protected-class baselines; and *Iqbal*-era pleading needs specific factual allegations---doctrinal predicates that anonymous transactions never generate. The effects fall asymmetrically: those most vulnerable to discrimination are least able to afford the shield and, when harms remain, least able to prove them. The challenge for the law shifts from detecting and remedying algorithmic discrimination to governing agent-mediated anonymity as civil rights infrastructure: ensuring access to privacy-preserving agents, regulating abuse without forced identification, and deciding whether retailers may refuse to deal with agents at all.

Anirban Mukherjee, H. Chang · 0 citations
Review Jul 2026

Private Again: AI Agents Restore Anonymity - Foreclosing Discrimination and Its Proof

Artificial intelligence agents can transact online on behalf of a human principal---browsing, paying, receiving, and reviewing---without revealing who that principal is. That architecture starves algorithmic discrimination of its inputs---identity, purchase history, location history, behavioral traces, and demographic proxies---but also forecloses its proof. Disparate-treatment needs comparators; disparate-impact needs protected-class baselines; and *Iqbal*-era pleading needs specific factual allegations---doctrinal predicates that anonymous transactions never generate. The effects fall asymmetrically: those most vulnerable to discrimination are least able to afford the shield and, when harms remain, least able to prove them. The challenge for the law shifts from detecting and remedying algorithmic discrimination to governing agent-mediated anonymity as civil rights infrastructure: ensuring access to privacy-preserving agents, regulating abuse without forced identification, and deciding whether retailers may refuse to deal with agents at all.

Anirban Mukherjee, H. Chang · 0 citations
Preprint Aug 2026

What survives honest evaluation? Leakage-safe, search-aware assessment of LLM-driven trading strategy discovery

Large language models (LLMs) are increasingly used to discover trading strategies, and much of the resulting literature shares a methodological weakness: many candidate strategies are generated, the best is reported, and neither look-ahead bias nor the intensity of the search behind the reported result is corrected for. We present a strategy-discovery system that makes both corrections structural rather than procedural. First, the agent can only act through registry-validated tools whose feature space excludes look-ahead by construction; we show that this guardrail is not redundant with statistical correction: a deliberately leaky oracle posting a Sharpe ratio of 35 survives Deflated Sharpe and probability-of-backtest-overfitting testing completely. Second, the system records every strategy evaluation its search performs and deflates all reported performance by that trial count, tracing how the best in-sample Sharpe ratio climbs with each trial while the deflation threshold, driven by the agent's own search, climbs faster. Across a 453-stock point-in-time US equity universe and a 39-ETF multi-asset universe with realistic transaction, impact, and borrow costs, honest evaluation certifies passive benchmarks (out-of-sample confidence intervals excluding zero), rejects every LLM-discovered strategy (across two frontier models, search budgets up to one hundred candidates, and five repeated runs), catching selection luck, predicted rank degradation, and out-of-sample collapse through complementary instruments, and evaluates a human trader's production rule system under identical instruments. The framework formalizes why pre-registered hypotheses earn lower evidential bars than brute search, and quantifies the sample sizes that credible certification of moderate edges actually requires.

Eray Gençay · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.