Skip to content
Open access

Automating Plan Evaluation Using Agentic Large Language Models

Sep 2026 · Journal of Planning Education and Research · Vol 46, pp. 699 - 712 · 1 citation · 42 references

TL;DR

This research benchmarks human evaluations against a large language model (LLM) using a multi-agent approach and/or retrieval-augmented generation (RAG) to automate complex content analysis tasks to leverage artificial intelligence’s efficiency and precision alongside humans’ contextual understanding and domain expertise.

Abstract

Manual plan evaluation faces reliability and scalability challenges. This research benchmarks human evaluations against a large language model (LLM) using a multi-agent approach and/or retrieval-augmented generation (RAG) to automate complex content analysis tasks. We find that LLMs generally perform comparably with humans, with most errors arising from overimplication and limited domain knowledge. The multi-agent approach substantially enhances LLM’s performance, reducing common machine errors by over 50 percent. Integrating such LLM tools with human oversight will likely become the new norm for content analysis, and this study demonstrates how to leverage artificial intelligence’s (AI) efficiency and precision alongside humans’ contextual understanding and domain expertise.

Read PDF

Similar papers

Open access Jul 2026

Language Model Council: A Multi-Agent Framework using Explainable AI

This work introduces the Language Model Council (LMC), a collaborative framework that combines the expertise of multiple specialized AI agents to evaluate a user query from different perspectives and outperforms traditional single-model systems by improving response quality, reducing hallucinations, and increasing user trust through enhanced explainability.

D. M, Shwetha Kr, G. Divya et al. · 0 citations
Preprint Aug 2026

From Atomic to Agentic: Towards Interpretable Evaluation of LLMs'Agentic Mathematical Capabilities

Experiments reveal that models with similar end-to-end accuracy can exhibit markedly different agentic capability profiles, demonstrating that process-level evaluation is crucial for interpreting the true potential of LLMs and guiding the development of next-generation mathematical agents.

Jiayi Kuang, Yinghui Li, Yun-Ze Song et al. · 0 citations
Review Open access Jul 2026

LLM-Powered Agentic Data Science: Automated Analysis and Insight Generation

It is argued that verification, not generation, is the binding constraint for trustworthy automated analysis in agentic data science: systems in which an LLM coordinates exploratory analysis, query generation, hypothesis formation, and reporting with limited human supervision.

M. Keerthika · 0 citations
Preprint Aug 2026

An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures

To test whether the taxonomy supports mitigation, TART, Taxonomy-Guided Actionable Representation, is introduced that makes the taxonomy's key aspects explicit to the planner and downstream sub-agents and consistently improves performance.

Vikas Pahuja, J. Brokman, O. Hofman et al. · 0 citations
Open access Aug 2026

STARS: extending interactive task learning with large language models

It is demonstrated that LLMs speed agent learning and greatly reduce the human effort required to achieve robust, reliable, and repeatable task performance.

James R. Kirk, Robert E. Wray, John E. Laird · 0 citations
Review

A Review on Test-Time Scaling for Agentic Large Language Models

A novel RAIE taxonomy along four scaling dimensions is proposed, which optimizes the entire thought process through search algorithms and self-verification, and introduces a task-oriented guideline for choosing the best TTS strategy.

Jia-Yu An, Zheng Chen, Yongcheng Jing et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.