Quality Score: A Behavioural Metric for Deliberation Quality in LLM-MAS Systems
Autonomous systems increasingly employ Large Language Models (LLMs) as a deliberative layer in situations where classical machine learning methods encounter out-of-distribution scenarios. The use of multiple models in a multi-agent configuration (LLM-MAS) enables mutual validation of responses and reduces the risk of e...