Skip to content

Reinforcement Learning of Communication in a Mesh of Small Language Models

Sep 2026 · 0 citations · 33 references
Computer Science

TL;DR

TalkMesh, a decentralized mesh of small language model agents that learns when and what to communicate is presented, a decentralized mesh of small language model agents that reaches the accuracy of majority voting over 32 samples with each of three models.

Abstract

Language models gain accuracy from more compute at test time, but majority voting over independent samples saturates: as samples grow, the vote converges to the model's most frequent answer. Communication can add what sampling cannot: an agent that solves a problem can pass the key step to the others. We present TalkMesh, a decentralized mesh of small language model agents that learns when and what to communicate. Each agent samples a proposal and scores it with a trained confidence head. The most confident agent broadcasts a hint; agents below a confidence threshold revise, keeping each revision that outscores its proposal. Gossip consensus approximates the vote weighted by confidence without a coordinator. A talk policy, trained with group relative policy optimization on the change in correctness after revision, writes hints and revisions. With three agents, which together generate at most six outputs, the mesh reaches the accuracy of majority voting over 32 samples with each of three models. Trained with at most 8 agents and evaluated with 32, it raises accuracy from 0.568 under self-consistency to 0.705 (Qwen3.5-0.8B, GSM8K) and from 0.492 to 0.722 (SmolLM3-3B, MATH-500). When 4 of 8 agents collude on a wrong answer with fabricated confidence and poisoned hints, majority vote accuracy falls to 0.000 (Qwen3.5-0.8B, GSM8K). A defended mesh, whose agents rescore solutions with their own confidence heads, retains 0.507. Across reasoning, embodied coordination, and traffic signal control, messages improve a decision when the acting agent cannot observe the information it requires and another agent can send it.

View source

Similar papers

#natural language process... Preprint Sep 2026

Message capacity and claim wording set the transition points of collective truth-finding in language-model networks

Whether human or large language model (LLM), an agent in a discussion reads only a few of the others'contributions, bounded by cognition, context, or cost. LLM collectives can settle on a wrong consensus even when a majority starts out correct; we ask how far that reading bound alone decides the outcome. We model the b...

Makoto Fukushima · 1 citation
#natural language process... Preprint Aug 2026

Benchmarking large language model agent societies against human behavioural distributions

SILICA is an open instrument that tests three doubts of large language model agents: whether the agents behave like the humans they stand in for, whether a finding survives changes to the apparatus that leave the rules untouched, and whether apparent social dynamics are interaction at all rather than the reproduction o...

Raad Bin Tareaf · 0 citations
#artificial intelligence Preprint Sep 2026

What Does Post-Training Change in Multilingual Reasoning?

These stages establish a constructive post-training path from English-pivoted capability to multilingual reasoning that is reliably delivered, and jointly track correctness, language adherence, termination, and delivery efficiency.

Hong-Yang Li, Xiao Li, Caesar Wu et al. · 0 citations
Preprint Aug 2026

Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation

A reasoning model is built that adaptively chooses how much to reason for each problem, and the brief modes end up more accurate than \textsc{Long}, which shows that the router sorts problems by difficulty rather than at random.

Gijs Kassenaar, Zhao Yang, Vincent François-Lavet · 1 citation
#machine learning Preprint Sep 2026

Self-Repulsive Sampling for Diffusion Language Models

Sampling several responses and voting over their answers can improve a language model's accuracy, but repeated answers limit the benefit of additional samples. Raising temperature increases diversity at a potential cost to per-sample accuracy. We introduce Self-Repulsion (SR), a sampler for masked diffusion language mo...

Michael A. Helcig, Martin Jaggi · 0 citations

Related blog posts

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

GPT-Lab Aug 28, 2026

We built an AI factory for HVAC control

What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.

Microsoft Research Blog Jul 30, 2026

EvoLib: Turning experience into evolving knowledge

LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib: Turning experience into evolving knowledge appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.