Jul 2026· Social science computer review· 0 citations· 48 references
TL;DR
Examination of the impact of agent diversity, agent open-mindedness, and human–AI collaboration (HAIC) in a multi-LLM-agent system for automated content analysis demonstrates reliable and accurate measurement of four communication variables across three datasets, with improved performance following agent discussion.
Abstract
Emerging research in computational social science has applied LLMs to automate content analysis, often by prompting a single model to act as a human coder. While a single LLM may suffice for a few manifest variables, it still falls short on diverse latent constructs. And the impact of LLM agent attributes on measurement outcomes remains unclear, limiting their validity for communication research. Drawing upon the literature on interacting agents and communication, this study examines the impact of agent diversity, agent open-mindedness, and human–AI collaboration (HAIC) in a multi-LLM-agent system for automated content analysis. The results demonstrate reliable and accurate measurement of four communication variables across three datasets, with improved performance following agent discussion. Additionally, agent open-mindedness, but not agent diversity, significantly affects measurement outcomes. These results highlight the potential of multi-LLM-agent systems for automated content analysis and suggest the importance of considering agent attributes and values in system design.
Identifying and promoting effective user interaction strategies in human-AI interaction is critical to improving collaboration quality. However, what motivates users or how agent behavior affects them remains unclear. We analyzed self-reported user strategies in a study (N=60) using a collaborative game with three AI partners reflecting leader, follower, and shifting initiative behavior. We identified four interaction dimensions — intention, perception, communication, and coordination — and assessed their relation to agent behavior, team performance, and user preferences. We observed that while users self-report goal prioritization independently of agent behavior, in-game logs reveal that non-collaborative goals are prioritized significantly more often with AI leaders than with followers. Moreover, users tend to construct plans early through observation when the AI leads, yet they are more likely to use strategies to propose goals when the AI follows. Our work provides a new framework for AI agent designers to understand user behavior and guidelines for better supporting it.
Inês Lobo, Janin Koch, Jennifer Renoux et al.· Proceedings of the 26th ACM...· 0 citations
With increasingly powerful large language models (LLMs) and LLM-based agents tackling an ever-growing list of tasks, we envision a future where numerous LLM agents work seamlessly with other AI agents and humans to solve complex problems and enhance daily life. To achieve these goals, LLM agents must develop collaborative skills such as effective persuasion, assertion and disagreement, which are often overlooked in the prevalent single-turn training and evaluation of LLMs. In this work, we present Collaborative Reasoner ( Coral ), a framework to evaluate and improve the collaborative reasoning abilities of language models. In particular, tasks and metrics in Coral necessitate agents to disagree with incorrect solutions, convince their partners of a correct solution, and ultimately agree as a team to commit to a final solution, all through a natural multi-turn conversation. Through comprehensive evaluation on six collaborative reasoning tasks covering domains of coding, math, scientific QA and social reasoning, we show that current models cannot effectively collaborate due to undesirable social behaviors, collapsing even on problems that they can solve singlehandedly. To improve the collaborative reasoning capabilities of LLMs, we propose a self-play method to generate synthetic multi-turn preference data and further train the language models to be better collaborators. Experiments with Llama-3.1 , Ministral and Qwen-2.5 models show that our proposed self-improvement approach consistently outperforms finetuned chain-of-thought performance of the same base model, yielding gains up to 16.7% absolute. Human evaluations show that the models exhibit more effective disagreement and produce more natural conversations after training on our synthetic interaction data. 1
Ansong Ni, Ruta Desai, Yang Li et al.· Neural Information Processin...· 7 citations
Firms often seek to replace human agents (HAs) with automated agents (AAs) for customer service, but customers still tend to react more negatively to AAs than to HAs. This meta‐analysis examines agent‐related perceptions (i.e., humanlikeness, warmth, and competence), which represent the key upstream customer reactions shaping downstream responses. Specifically, it clarifies the hierarchy of these perceptions and examines the moderating effect of task‐related risks (i.e., functional, psychological, and social) on these perceptions. Meta‐analytic path modeling with subgroup analyses, including 1191 effect sizes, provides novel insights. First, there is indication that perceived humanlikeness is not a prerequisite for perceived warmth and competence; all three perceptions emerge in parallel. Second, tasks entailing functional or psychological risk largely weaken the negative perceptions of AAs (vs. HAs). For psychologically risky tasks, AAs are not perceived as less competent than HAs. Tasks entailing social risk largely strengthen the negative perceptions of AAs (vs. HAs). These findings challenge two prevailing assumptions in the services marketing literature: that customers favor humans over AAs simply because they are not humanlike and that risk generally discourages AA usage. The results have important managerial and future research implications.
Sandra Miederer, Katja Gelbrich, H. Roschk et al.· Psychology & Marketing· 0 citations
An automated, data-driven approach to uncover patterns, which the authors term traits, of effective human-AI interaction that are aligned with task outcomes is explored and Principal Trait Analysis is proposed, a Principal Component Analysis-inspired algorithm for deriving common traits from patterns in LLM conversations.
Hunter McNichols, Kai Du, Andrew S. Lan· 0 citations
This work presents AgentPanel, a multi-agent forum for human--AI collaboration in scientific exploration, a multi-agent forum for human--AI collaboration in scientific exploration that outperforms a centralized multi-agent debate baseline and shows that users value AgentPanel for perspective diversity and exploration support.
Zhiyao Cui, Qianyi Wang, Hao Yan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.