Skip to content

CALICO: A Human-Centered, Codebook-Aligned System for Annotation

Sep 2026 · 0 citations · 35 references
Computer Science

TL;DR

CALICO is presented, a human-centered, codebook-aligned annotation workflow that treats prompts as editable, versioned, and optimizable artifacts and integrates codebook parsing, prompt generation, result inspection, prompt versioning, natural language human feedback, and label-supervised prompt optimization through existing optimizers.

Abstract

Large language models are increasingly used to scale codebook-based annotation in scientific research, but existing workflows provide limited support for translating domain experts'codebooks into reliable, revisable, and auditable prompts. Prompts are often treated as fixed instructions and hidden from annotators, making it difficult for non-technical domain experts to diagnose and correct model behavior when outputs violate codebook guidelines. In this paper, we present CALICO, a human-centered, codebook-aligned annotation workflow that treats prompts as editable, versioned, and optimizable artifacts. CALICO integrates codebook parsing, prompt generation, result inspection, prompt versioning, natural language human feedback, and label-supervised prompt optimization through existing optimizers such as GEPA, MIPROv2, and OPRO, together with our reflection-based optimizer, ReflectAgent. Empirically, we evaluate CALICO on domain-specific AI-companion chatbot conversation codebooks. Across evaluated dimensions, CALICO improves mean held-out performance by +13.0 and +7.4 absolute points for two coders, respectively. A coder-specificity analysis further suggests that optimized prompts capture coder-specific interpretations rather than only generic codebook clarification. CALICO runs as a web application that takes users from raw codebook materials to inspectable, exportable labels; the website, codebase, and live demo are released at https://calico-annotation.github.io/ under the Apache 2.0 License.

View source

Similar papers

Open access Sep 2026

Using Large Language Models for Automated Corpus Annotation and Linguistic Analysis: A Critical Methodological Framework

Large language models (LLMs) are increasingly used to classify, label, summarize, and interpret large text collections, creating new possibilities for corpus linguistics. Their capacity for zero-shot and few-shot instruction following could reduce the cost of linguistic annotation and extend analysis beyond the categor...

Maria Ibrar · 0 citations
Preprint Aug 2026

Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization

This work presents AudioChaps, a post-training framework for aligning end-to-end LALMs for this task via Group Relative Policy Optimization (GRPO) guided by Chain-of-Thought (CoT) reasoning, and demonstrates that GRPO-trained LALMs can reliably transform unstructured auditory streams into navigable, structured media.

Tony Alex, Wish Suharitdamrong, Sara Atito et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation

Large-scale text annotation brings expert insight to millions of documents, often through a codebook that AI annotators follow. Developing a robust codebook, however, takes months. Large language models (LLMs) could speed this process by applying an early codebook to the data, surfacing cases with strong LLM disagreeme...

Ze-Yu He, Zhu-Qian Zhou, Kirk P. Vanacore et al. · 0 citations
Review Sep 2026

VISTA: Dense Multi-Label Classroom Coding with Vision-Language Models

Video-language benchmarks are usually constructed by the dataset authors without published reliability statistics, leaving the noise floor of the construct unknown. We argue that multimodal benchmarking benefits from methods taken from research communities that have already invested in strategies to ensure reliability....

Andrew Franck, Brendan Ng, Ben Fitzgerald et al. · 0 citations
Review Aug 2026

AnchorScore: A CLIP-Based Diagnostic of MLLM Annotation Difficulty

AnchorScore provides a low-cost ranking signal that directs expensive MLLM evaluation to the classes where it is most informative, and three practical applications follow: a deployable hybrid CLIP/MLLM routing strategy, prompt disambiguation on hard classes (exploratory), and review-priority prediction for human verifi...

Yanzhang Ma, Li-Zhuo Zhang · 2 citations
Preprint Aug 2026

ChartAnno: Benchmarking Multimodal Large Language Models for Chart Annotation Generation

This work introduces ChartAnno, a comprehensive benchmark for evaluating MLLMs on chart annotation generation, and develops a multidimensional evaluation framework combining rule-based and LLM-judged metrics to assess execution, structural compliance, semantic consistency, and design effectiveness.

Zhenghan Chen, Zekai Shao, Lidan Tan et al. · 0 citations

Related blog posts

Microsoft Research Blog Jul 8, 2026

Flint: A visualization language for the AI era

Short chart specifications are easy to write, but often produce uninspiring results. Flint is an open-source visualization language that offers a middle path, letting AI agents create expressive charts from compact, human-editable specifications. The post Flint: A visualization language for the AI era appeared first on Microsoft Research.

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.