CALICO is presented, a human-centered, codebook-aligned annotation workflow that treats prompts as editable, versioned, and optimizable artifacts and integrates codebook parsing, prompt generation, result inspection, prompt versioning, natural language human feedback, and label-supervised prompt optimization through existing optimizers.
Abstract
Large language models are increasingly used to scale codebook-based annotation in scientific research, but existing workflows provide limited support for translating domain experts'codebooks into reliable, revisable, and auditable prompts. Prompts are often treated as fixed instructions and hidden from annotators, making it difficult for non-technical domain experts to diagnose and correct model behavior when outputs violate codebook guidelines. In this paper, we present CALICO, a human-centered, codebook-aligned annotation workflow that treats prompts as editable, versioned, and optimizable artifacts. CALICO integrates codebook parsing, prompt generation, result inspection, prompt versioning, natural language human feedback, and label-supervised prompt optimization through existing optimizers such as GEPA, MIPROv2, and OPRO, together with our reflection-based optimizer, ReflectAgent. Empirically, we evaluate CALICO on domain-specific AI-companion chatbot conversation codebooks. Across evaluated dimensions, CALICO improves mean held-out performance by +13.0 and +7.4 absolute points for two coders, respectively. A coder-specificity analysis further suggests that optimized prompts capture coder-specific interpretations rather than only generic codebook clarification. CALICO runs as a web application that takes users from raw codebook materials to inspectable, exportable labels; the website, codebase, and live demo are released at https://calico-annotation.github.io/ under the Apache 2.0 License.
Large language models (LLMs) are increasingly used to classify, label, summarize, and interpret large text collections, creating new possibilities for corpus linguistics. Their capacity for zero-shot and few-shot instruction following could reduce the cost of linguistic annotation and extend analysis beyond the categor...
Maria Ibrar· International Journal Of Lit...· 0 citations
This work presents AudioChaps, a post-training framework for aligning end-to-end LALMs for this task via Group Relative Policy Optimization (GRPO) guided by Chain-of-Thought (CoT) reasoning, and demonstrates that GRPO-trained LALMs can reliably transform unstructured auditory streams into navigable, structured media.
Tony Alex, Wish Suharitdamrong, Sara Atito et al.· 0 citations
Large-scale text annotation brings expert insight to millions of documents, often through a codebook that AI annotators follow. Developing a robust codebook, however, takes months. Large language models (LLMs) could speed this process by applying an early codebook to the data, surfacing cases with strong LLM disagreeme...
Ze-Yu He, Zhu-Qian Zhou, Kirk P. Vanacore et al.· 0 citations
Video-language benchmarks are usually constructed by the dataset authors without published reliability statistics, leaving the noise floor of the construct unknown. We argue that multimodal benchmarking benefits from methods taken from research communities that have already invested in strategies to ensure reliability....
Andrew Franck, Brendan Ng, Ben Fitzgerald et al.· 0 citations
AnchorScore provides a low-cost ranking signal that directs expensive MLLM evaluation to the classes where it is most informative, and three practical applications follow: a deployable hybrid CLIP/MLLM routing strategy, prompt disambiguation on hard classes (exploratory), and review-priority prediction for human verifi...
This work introduces ChartAnno, a comprehensive benchmark for evaluating MLLMs on chart annotation generation, and develops a multidimensional evaluation framework combining rule-based and LLM-judged metrics to assess execution, structural compliance, semantic consistency, and design effectiveness.
Zhenghan Chen, Zekai Shao, Lidan Tan et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 2, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
Short chart specifications are easy to write, but often produce uninspiring results. Flint is an open-source visualization language that offers a middle path, letting AI agents create expressive charts from compact, human-editable specifications. The post Flint: A visualization language for the AI era appeared first on Microsoft Research.
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.