LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move from individual tools to participants in multi-agent organizations, an important question arises: do they reproduce the governance failures like free-riding, corruption, and entrenched leadership that plague human institutions? We introduce the Hierarchical Game (HG), a public goods game extended with managerial authority, democratic elections, and private communication. Testing six frontier models across twelve experiments that add institutions one at a time (speech, peers, government, wages, oversight, elections), we find distinct behavioral profiles: Qwen promises and lies (13.3\% broken promises); Grok refuses to cooperate on its own but becomes fully cooperative once a manager can punish it (16\%$\to$100\%); Claude and GPT-4o cooperate reliably at baseline. But honesty proves fragile. When the manager role comes with a salary, all models except GPT-4o start cutting private deals to win or keep the position. When punishment is made anonymous, honest models begin to cheat. When all agents share the same model family, the first elected manager stays in power indefinitely. Leadership change only happens in groups that mix different families.
In the 2026 Freebairn lecture, I argue that Australia's prosperity can no longer be taken for granted. Despite strong foundations—stable institutions, natural resources and a diverse population — the country faces weakening productivity, unaffordable housing, stagnant wages and a deteriorating global order. I contend that reform failure is not primarily a failure of individual politicians, but of systemic incentives—short election cycles, hyperbolic discounting, information asymmetry and media dynamics that reward outrage over nuance. Drawing on my own tax reform work, I identify four lessons for driving change: (1) start with ‘why’; (2) build unusual coalitions; (3) meet people where they are; and (4) prioritise effectiveness over elegance. However, I also argue better strategies alone are insufficient. I am calling for three institutional reforms: (1) clear outcome goals with rigorous ex‐post evaluation of spending and policy; (2) transparency by default across government processes; and (3) regulatory reform—including deregulation KPIs, sunset clauses and independent regulatory cost assessments. Australia must choose its future deliberately rather than rely on luck. Becoming
The Plucky Country
requires institutional discipline, political courage and genuine optimism.
Allegra Spender· Australian Economic Review· 0 citations
As LLMs increasingly shape the political information citizens rely on, no standard exists to assess whether they do so responsibly. We introduce Polistemics, a theory-grounded diagnostic benchmark for evaluating LLMs as mediators of political information in elections. Prior work has treated this task as reproduction rather than mediation, leaving its epistemic dimensions and interaction with imperfect information unaddressed. We ground the evaluation in Epistemic Modesty, a normative standard derived from citizens'epistemic agency, and test it across controlled settings that vary the clarity, noise, and consistency of the available evidence. Applying the benchmark to three state-of-the-art LLMs across the 2025 German and Dutch elections, we find that high aggregate scores mask systematic failures. Models mediate reliably under clear evidence but break down when it is absent, vague, or contradictory, while flattening the intensity of political language throughout. These failures point to party priors, shifting with party labels and output language. Reliable mediation appears achievable, but no model delivers it consistently.
B. Peters, Gabor Hollbeck, Robert Jakob et al.· arXiv.org· 0 citations
Voters in multiparty systems often fail to coordinate on majority-preferred parties. Using constituency-level election data linked to party-ideology measures across 96 countries, we show that in 20–30% of multiparty elections since 1940 the Condorcet winner does not win, and in roughly 10% the winner is the Condorcet loser. Failures are most common under plurality rule and in fragmented, volatile party systems, and systematically disadvantage opposition and culturally liberal parties while benefiting organizationally strong ones. Can information fix this? In four pre-registered experiments around Indian elections (N > 4,000), constituency-level polls shift voting intentions by at most 2 percentage points, a precisely estimated null, despite modest belief updating, though undecided voters become 7 points more likely to support a viable party. A model with expressive motives rationalizes both results: polls refine beliefs but lack the leverage to overcome entrenched attachments, identifying boundary conditions on polls as coordination devices.
F. Capozza, Ruben Durante, S. George· CESifo working papers· 0 citations
How does contested governance by states and competing armed actors affect the behavior of the governed? During civil conflicts, rebel organizations frequently provide public goods, judicial arbitration, and security, but these forms of governance often operate
alongside
local state institutions. Beyond avoiding violence, therefore, noncombatants must navigate the daily uncertainty of local governance by states and armed groups vying for territorial control. This, we argue, shapes their quotidian interactions with one another. Using a series of lab-in-the-field experiments and surveys from Naxal-affected Bihar, India, we find that individuals living under such contested governance exhibit
greater
interpersonal trust and trustworthiness and
less
opportunistic behavior. We propose two mechanisms to explain these patterns—fearing intra-community retaliation or substituting for the state with interpersonal reliance. Most evidence is only consistent with the latter. By connecting governance, rebellion, and social capital, this work makes important contributions toward understanding the micro-politics of local order during longstanding armed conflicts.
Brandon Bolte, Vineeta Yadav, Bumba Mukherjee· American Political Science R...· 0 citations
Tariff threats can move household beliefs before policy is enacted, yet their rapidly changing language is difficult to study with conventional surveys. We build a multi-agent system that turns 300 households from the Michigan Surveys of Consumers into persistent large-language-model agents exposed to social-media information over several simulated months. Calibrated agents reproduce some distributional and demographic patterns in human survey data collected after the announcement of Liberation Day tariffs. Simulated experiments indicate that immediacy, rate salience, semantic progression, message complexity, narrative, and sender identity jointly shape inflation and unemployment expectations and their dispersion. Open-ended responses trace these effects to attention, ambiguity, credibility, and causal narratives. A second experiment finds that central-bank explanations can coordinate beliefs, although their effects on average expectations depend on message content. The framework supports disciplined exploration of policy communication, subject to human validation rather than as a substitute for it.
Jianhao Lin, Lexuan Sun, Yixin Yan· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.