Skip to content

RFCLLM: Evaluating LLMs'Reasoning Ability of Network Protocol State Machines

Sep 2026 · 0 citations · 33 references
Computer Science

TL;DR

This work examines the degree to which an LLM's implicit representation of a finite-state transition system-defined via natural language descriptions-aligns with a manually generated ground-truth model.

Abstract

Mapping textual specifications into formal representations is essential for ensuring the correctness of protocol designs and implementations. LLM-generated mappings, used for networking security or testing, are assumed to capture a perfect understanding of the specification, which may not hold in practice. The goal of this paper is to assess the extent to which LLMs can interpret the specification correctly. We examine the degree to which an LLM's implicit representation of a finite-state transition system-defined via natural language descriptions-aligns with a manually generated ground-truth model. We designed 4 tasks and 1482 task queries for 16 protocols. We evaluated different judge biases, observed the inherent difficulty gaps between tasks, looked into the effect of 4 context types, and the influence of protocol characteristics. Our work contributes to a step toward verifying whether LLMs can really be trusted in FSM (Finite State Machine) reasoning of protocol specifications.

View source

Similar papers

Preprint Sep 2026

Can LLMs help find Ambiguities in Protocol Specifications?

Internet protocol specifications written in RFCs are subject to ambiguities and multiple interpretations that can cause interoperability failure. While these have presumably cleared up after years of experience, such ambiguities can bedevil the adoption of newer protocols like 5G. The 5G specifications pair a formal me...

Zi-Yue Dang, Si-Xu Tan, Atharva Nevasekar et al. · 0 citations
Preprint Sep 2026

Confidence-Guided Protocol IR for LLM-Aided Security Protocol Modeling

Large language models offer a promising interface for translating natural-language protocol descriptions into formal security models, but their outputs remain difficult to trust without expert validation. In this paper, we present a human-in-the-loop framework for generating Tamarin-verifiable formal models of security...

Si-Qi Li, Yu-Fan Cai, Hong-Shu Wang et al. · 0 citations
Preprint Sep 2026

NetInspector: Measuring and Improving LLM Capabilities for Reliable Intent-Based Networking Policy Generation

Modern networks are large in scale and heterogeneous in configuration, making manual policy management increasingly impractical. Intent-Based Networking (IBN) addresses this by automating the translation of high-level operator goals into low-level network configurations. Yet existing IBN systems rely on static heuristi...

Yuxuan Zhang, Hongxin Hu, Guo-Fei Gu · 1 citation
Preprint Aug 2026

NoTB: Oracle-Free Triage of LLM-Generated RTL via Cross-Model Formal Consensus

NoTB is introduced, an oracle-free triage framework that infers correctness from cross-model formal consensus and demonstrates that formal cross-model agreement provides a reliable basis for high-confidence triage without model-dependent oracles.

Elisavet Lydia Alvanaki, Je Yang, Biruk B. Seyoum et al. · 0 citations
Preprint Sep 2026

LLM-Assisted Automatic Security Proofs for Cryptographic Protocols: How Far Are We?

Large language models (LLMs) have shown strong potential for assisting software and security analysis tasks, yet their effectiveness in cryptographic symbolic protocol verification remains insufficiently understood. In this paper, we conduct the first systematic evaluation of the capability of state-of-the-art LLMs in...

Tian-Jian Liu, Shi-Cheng Feng, Jin'ao Shang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.