Skip to content
Review Open access

Workflow Signal Protocol: A Measurement Method for Deployment-Time Workflow Observability in Legal AI

2026 · IEEE Access · Vol 14, pp. 116439-116457 · 0 citations · 62 references
Computer Science

TL;DR

The Workflow Signal Protocol (WSP) is introduced, a deployment-layer measurement method for recording workflow observability as a structured workflow-observability record that supports the central claim that deployment-time workflow observability is measurable within the evaluated legal-AI settings.

Abstract

Legal AI benchmarks, citation checks, and retrieval-grounding tests primarily evaluate upstream capability: whether a model can answer, extract, or ground a legal task. Deployment asks a different question: whether a particular output remains observable enough to be deployed, reviewed, corrected, or escalated once it enters an organizational workflow. We introduce the Workflow Signal Protocol (WSP), a deployment-layer measurement method for recording workflow observability as a structured workflow-observability record. WSP encodes source status, proposition support, review state, recourse, provenance, and role-scoped disclosure. We validate WSP through controlled stress tests, public legal datasets, documented real-world failures, and a live-output pilot using three general-purpose model application programming interface (API) arms. In the main matched-vocabulary stress test, local formal/substantive routing reduced hidden-risk deployment from 93.5% under calibration-only abstention to 4.0% or below; all 192 pilot outputs were expressible as WSP records. These results support the central claim that deployment-time workflow observability is measurable within the evaluated legal-AI settings; validation in deployed institutional legal workflows remains future work.

Read PDF

Similar papers

#artificial intelligence Review Aug 2026

LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents

Large language model (LLM) agents can now carry out long-horizon technical workflows involving complex tool use, code execution, file edits, and generated artifacts. As agents do more work faster, the productivity bottleneck shifts from producing outputs to auditing whether those outputs are correct and trustworthy. Agent observability systems make fine-grained execution events visible, but visibility alone still leaves reviewers to reconstruct which actions, artifacts, and validation steps matter for a particular conclusion. We introduce LEDGER - Layered Evidence and Decision Graphs for Execution Review, a tracing and review system that builds layered trace graphs over observed agent sessions. LEDGER preserves Trace Records while grouping them into Evidence Nodes and Workflow Nodes, representing artifacts as evidence anchors, and adding typed semantic edges that connect claims to supporting actions, artifacts, and checks. Through data-analysis and coding examples, we show how the resulting traces expose workflow decisions, artifact lineage, repair steps, validation coverage, and claim-support paths for evidence-centered audit.

Daehong Kim, Hai-Chao Miao, Shusen Liu · 2 citations
#artificial intelligence Preprint Sep 2026

IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier

Enterprises deploy systems, not checkpoints. Usable capability depends jointly on weights, serving route, precision, output contract, and harness, yet all 18 audited benchmarks score advertised model identifiers. We treat this as measurement error and give a protocol that makes it reportable. It has three parts. A gold-blind capability-binding preflight verifies that a route can execute the evaluation contract before any task reaches it; a reliability-inclusive first-pass scoring rule keeps failure in the score while keeping unsupported capability out; and adjudication is structurally score-blind. We call the protocol IB2 and release its algorithms, classification tables, request contract, and manifest schemas. Its reference instantiation, 128 locked tasks and 987 assertions over document, spreadsheet, chart, tool and database work, stays sealed: the procedure is the artifact, not the corpus. Across eleven systems, four results. Capability availability is measurable: two complete single-route runs on identical weights later failed distinct predicates of the finalized binding gate, while a third passed that gate before a fresh run. The advertised identifier exposed neither limit. Discrimination is not uniform: four of seven suites saturate under a six-system band, with the spread almost entirely from governed database work and multi-tab joins, so we report interval-backed resolution groups, not ranks; two of the nominal five-label output's four cuts fail multiplicity adjustment. Serving-arm choice moved one declared revision and precision from 77.38 to 82.54, paired interval [0.11,10.60], though the arms differ in access mode, harness generation, and the serving tool-call parser, and harness generation is a property of our evaluator, not any endpoint. Excluding failed responses from denominators changes the point ordering, so reliability inclusion changes a conclusion, not its wording.

Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan · 0 citations
Open access Jul 2026

Contextualized, Trustworthy, and Collective Scientific Decision Workflows

Science gateways often capture workflow execution but not the rationale behind human decisions such as dataset selection, algorithm choice, parameterization, and result interpretation, which limits reproducibility. We propose a framework that captures and shares scientific decisions as verifiable and reusable knowledge assets. We formalize decision workflows as a Markov Decision Process (MDP), in which experimental contexts are defined as states, user choices as actions, and multi-criteria constraints as reward functions. To ensure tamper-evident provenance and secure attribution, decision records use JSON-LD linked to Decentralized Identifiers as Verifiable Credentials (VCs). The architecture consists of three layers: a User layer for interaction and control, a Service layer for ranking and credential management, and a Data layer for semantic storage and blockchain-based verification. We implemented a prototype on the Ethereum Sepolia testnet to demonstrate end-to-end decision capture, credential issuance, selective disclosure, and blockchain-based verification. In a user study, all finalized decisions were captured as structured records, and 45 of 51 were issued as user-controlled VCs. The MDP-based ranking achieved a Top-1 agreement of 92.2% with user selections, substantially outperforming an approximate random baseline. Interaction patterns suggest two modes: guided exploration for non-expert users and selective deviation from ranked recommendations by expert users. These results indicate that decision provenance can complement traditional workflow provenance by preserving the reasoning behind scientific choices. By combining decentralized identity, semantic interoperability, and probabilistic decision modeling, the framework supports transparent, user-controlled, and collaborative experimentation and provides a basis for collective reuse of decision knowledge.

Pouriya Miri, V. Stankovski, D. Lavbič et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.