Skip to content
Preprint

LatticeMind: A Conflict-Aware Memory Primitive for Multi-Agent Systems

Aug 2026 · 0 citations · 45 references
Computer Science

TL;DR

LatticeMind is presented, a conflict-aware structured memory that handles contradiction at write time, which maintains explicit item status, applies cheap symbolic conflict checks, and invokes LLM reconciliation only for unresolved semantic cases.

Abstract

Multi-agent LLM systems often fail not for lack of candidate answers, but because they have no persistent mechanism for deciding which incompatible claim should currently be trusted. Majority vote, debate, and judge-based selection choose an output without recording which claim wins, which is contested, or why a later update supersedes it. We present \term{LatticeMind}, a conflict-aware structured memory that handles contradiction at write time. It maintains explicit item status, applies cheap symbolic conflict checks, and invokes LLM reconciliation only for unresolved semantic cases. On a label-blind ConflictBank evaluation that removes source-name hints, LatticeMind reaches 0.97 accuracy versus 0.61 for the strongest aggregation baseline, with the gap significant at $p<10^{-6}$ by paired McNemar test. Ablations show that removing the checker or the reconciler costs 12 to 14 points. On four secondary planning benchmarks the picture is mixed: LatticeMind beats naive merge on three of four, but does not replace deliberation methods on tasks rewarding iterative search.

View source

Similar papers

Preprint Sep 2026

Where Do Multi-Agent Systems Fail? Evidence-Grounded Diagnosis of Collective Mechanisms

When a multi-agent system answers correctly, it is tempting to conclude that its agents shared, checked, and used information as intended. Yet a system can break one of its collective mechanisms, the rules that govern how agents route, admit, store, and act on shared information, and still return the right answer, whil...

Zheng Han · 0 citations
Preprint Aug 2026

When Personal Memory Has No Single Answer: Evaluating LLM Agents under Irreducible Conflict

LLM agents increasingly maintain personal memory across sessions, but it can conflict. Preferences depend on context, behavior evolves, and sources can conflict. When a query lacks context, time, or source authority to interpret conflict, treating one memory as definitive converts unresolved conflict into an unjustifie...

Lulu Yang, Shusheng Xu, Zhuo-Ran Li et al. · 1 citation
#natural language process... Preprint Oct 2026

Right Answers, Wrong States: Hidden Information Failures in Multi-Agent Collaboration

Multi-agent systems are often judged by whether they reach the correct answer. This can miss a distinct failure: collaboration may leave behind a corrupted information state even when the immediate decision is correct. We call this an off-query failure. To study this failure in collaborative decision support, we introd...

He-Run Wan, Jia-Ying Wu, Min-Nan Luo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

Evaluating claim admission in shared agent memory is challenging because repeated claims may be mistaken for independent evidence. An agent may copy or paraphrase a retrieved belief, while admitting a false claim exposes subsequent agents to it. To study this problem, we introduce the Correlated Promotion Benchmark (CP...

Xiao-Yang Li, Yi-Qi Wang, Chen-Cheng Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DolphinBench: Mapping the Pareto Frontier of Agent Memory

DolphinBench is presented, a benchmark that evaluates memory directly through an agent's task completion and requires all evaluations to report total cost and latency alongside accuracy, which enables us to evaluate agent memory systems holistically.

Soumil Rathi, Deshraj Yadav, Taranjeet Singh · 0 citations
Preprint Aug 2026

Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages

Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages should shape a final answer. Such filtering assumes that a message likely to be correct is also worth keeping. Yet a wrong answer can contain a useful decomposition, constraint, or scientific principle. We test t...

Chih-Hsuan Yang, Anjir Ahmed Chowdhury, Cheng-Hau Yang et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.