The rapid integration of large language model-based conversational systems has intensified longstanding questions in mediated communication, human–computer interaction, and social presence research by placing language-generating systems within ongoing conversational exchanges. Rather than treating these systems only as information-retrieval tools, users often orient to them through interactional cues associated with responsiveness, continuity, and interlocutor recognition. However, existing analytical approaches do not adequately capture the multidimensional nature of these interactions. This study addresses this gap by proposing a computational framework grounded in communication theory, operationalized through three interaction indices: the Social Presence Index (SPI), the Social Bonding Index (SBI), and the Companion Communication Index (CCI). The framework is evaluated on more than 50,000 conversations from heterogeneous datasets (WildChat, LMSYS, DailyDialog, and MultiWOZ), enabling a comparative analysis between human–AI and human–human communication. Results show that human–human interactions exhibit higher social presence (SPI = 0.31 vs. 0.13), while human–AI interactions demonstrate greater structural continuity (CCI = 0.41 vs. 0.34), with statistically significant differences (
p
< 0.001). Segmentation analysis reveals non-linear interaction patterns, where conversational continuity peaks at intermediate levels of social presence in human–AI exchanges. These findings demonstrate that social presence, affective bonding, and conversational continuity are interdependent and context-sensitive, supporting a multidimensional framework for understanding and designing human–AI communication processes.
W. Villegas-Ch., Iván Ortiz-Garcés, Fernando Zúñiga-Tello et al.· Frontiers of Computer Scienc...· 0 citations
Large language models (LLMs) are increasingly used to evaluate open-ended educational responses. However, their performance is often assessed using aggregate metrics that provide limited insight into prediction stability, uncertainty, error patterns, and feedback quality. This study presents EduFairBench, a reproducible evaluation protocol designed to characterize LLM behavior across short-answer assessment and automated essay scoring using open educational benchmarks. The protocol combines repeated inference, majority-vote consolidation, uncertainty estimation, error analysis, and structural evaluation of generated feedback within a unified experimental framework. Experiments were conducted on SciEntsBank, Beetle, and ASAP2, comprising 2,000 student responses and 10,000 independent LLM inferences. The results showed moderate predictive agreement with human assessment while revealing substantial differences between nominal and ordinal evaluation tasks. Repeated inference demonstrated high internal stability across benchmarks, although systematic errors remained in semantically adjacent categories, indicating that prediction consistency does not necessarily imply correctness. Feedback quality varied by task type, with longer textual contexts yielding more specific and pedagogically structured explanations. These findings demonstrate that evaluating educational LLMs requires complementary analyses beyond conventional performance metrics. EduFairBench provides a reproducible methodology for jointly analyzing predictive performance, robustness, uncertainty, and feedback quality, providing a comprehensive methodological framework for the rigorous evaluation of LLM-based educational assessment systems.
W. Villegas-Ch., Aracely Mera-Navarrete, Fernando Zúñiga-Tello et al.· Frontiers in Artificial Inte...· 0 citations
This work reframes cross-model phishing detection from a problem of model incompatibility to one of practical calibration, and provides two deployable solutions, threshold recalibration on a small target slice and aggregated-pool training, along with a publicly released multi-LLM corpus.
Rommel Gutierrez, W. Villegas-Ch., Jaime Govea· Frontiers in Big Data· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.