This work characterize the least-cost separating pledge schedule, show that client stakes shift engagement toward or away from the least-skilled worker depending on the size of the pledge, and derive a stakes threshold above which building skill dominates free-riding on a rival's training - reversing the asymmetric-specialization result of the underlying labor model.
Abstract
Firms that deploy improving but imperfect AI must decide how much to keep human workers engaged. Engagement lowers current output yet builds the fallback skill the firm needs when AI fails. We ask what that fallback skill signals to two audiences at once: mobile workers, who sort across firms on the skill trajectory a job builds, and clients, who cannot observe skill in a credence-good market and must infer competence. We embed the engagement-skill dynamics of Singh et al. (2026) in a signaling game and add a liability commitment. Because a more-skilled provider fails less often precisely in the states where AI is down, the expected cost of a liability pledge is decreasing in fallback skill. This restores Spence-Mirrlees single crossing on a type that is endogenous - built, not drawn - and yields a separating equilibrium in which liability certifies preserved human competence that no artifact can certify once AI writes as well as the expert. We characterize the least-cost separating pledge schedule, show that client stakes shift engagement toward or away from the least-skilled worker depending on the size of the pledge, and derive a stakes threshold above which building skill dominates free-riding on a rival's training - reversing the asymmetric-specialization result of the underlying labor model. Two boundaries close the market from both sides: small tickets cannot fund enforcement, and large tickets exceed the provider's solvency. An agent-based version of the market reproduces the analytical thresholds under noisy beliefs, learning-by-record and worker churn.
It is shown that provenance certification priced as a type-independent stamp (e.g., C2PA) cannot restore full separation, while a verified commitment to forgo the AI frontier re-imposes the pre-AI artifact cost function.
Problem definition. When generative AI produces expert artifacts clients cannot distinguish from a competent provider's, the classical cost-based quality signal collapses and only outcome-contingent commitments can separate types. Such a commitment certifies an endogenous, perishable asset: the human fallback capabilit...
Lifelong LLM agents increasingly rely on external skill artifacts as one element for preserving and reusing capabilities over time. These skills (usually portable Markdown files such as SKILL.md) describe when and how to apply a capability and must be corrected, expanded, and consolidated as tools and usage patterns sh...
Empirical measures of AI exposure ask language models to score O*NET tasks for technical feasibility. In finance, technically feasible tasks must still pass through review, documentation, supervision, confidentiality controls, and accountable human sign-off before entering production. We measure the gap between feasibi...
Claes Backman, Christos A. Makridis· CESifo working papers· 0 citations
An AI that can only give advice seems safe: the human is always free to ignore it. That is the premise of the boxing tradition in AI safety, and its long-suspected weak point is that the human who reads the answers is part of the system. We make the fraction $\varepsilon_t$ of behavior that follows the advice a state o...
A common reassurance in AI safety holds that a system with benign terminal goals will behave accordingly. We argue that this reassurance fails structurally, and we identify where. For a capable agent that holds its objective as settled, a sense covering execution competence as well as content, continued human oversight...
A. Clark· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.