Back to feed
Conference

CAPTAIMN: A Real-Time LLM and RAG Based Decision Support System for Navigational Safety and COLREGs Compliance

Jul 2026 · International Conferences on Human-Machine Systems · pp. 622-627 · 0 citations · 31 references

Abstract

The correct regulatory interpretation in naval environments is challenging due to the complexity and urgency of decisions based on the International Regulations for Preventing Collisions at Sea (COLREGs). This article presents the development of an intelligent agent named Cognitive Agent for Analysis of Interrelated Problems in Maritime Navigation, hereafter referred to as CAPTAIMN. This agent integrates Large Language Models (LLM) and a Retrieval-Augmented Generation (RAG) architecture to support human decision-making and officer training in safety-critical naval environments. Thus, this work aims to propose a methodology for building an intelligent agent based on LLM and RAG, specifically focused on the assisted and contextualized interpretation of COLREGs. The proposed methodology was evaluated through a quantitative study with 15 maneuvering officers, who assessed 150 responses generated by a local language model using a Likert scale. The results from this phase showed significant approval, with 80% of the responses being rated as 'Agree' or 'Totally Agree' by the officers. These results suggest that the integration of LLM and RAG through CAPTAIMN can provide useful support for both decision-making and tactical training in naval operations.

View source

Similar papers

Preprint Aug 2026

Exploring LLM Capabilities for Situational Understanding and COLREG compliance on real-world maritime navigation scenarios

Recently, Large Language Models (LLMs) have shown considerable capability for situational understanding, reasoning, and decision making in different domains, most notable in the automotive sector. Therefore, we explore current state-of-the-art LLMs as a tool for maritime navigation, which includes both codified rules in the Collision Regulations (COLREGs) and uncodified best practices summarized in the concept of ``Good Seamanship''. We construct a dataset consisting of 50 diverse, real-world navigation scenarios from AIS data, label scenarios with applicable COLREG rules, recommended actions, and the reasoning for the action. We explore a variety of different LLM architectures and sizes to determine their understanding of maritime navigation tasks as well as evaluate their reasoning capabilities in this domain. The results obtained indicate that the maritime navigation task remains difficult to solve without fine-tuning, even for larger online models.

Julius Wirbel, P. N. Hansen, Line Clemmensen et al. · 0 citations
Open access Jul 2026

Transforming Maritime SAR Operations: Towards a Theoretical Framework for Human-AI Collaboration

This paper examines how Generative AI and Large Language Models (LLMs) can support decision-making in Maritime Search and Rescue (MSAR) operations through a human-centered collaborative framework. The authors propose a theoretical Human-AI Deliberative Framework (HADF) based on the ExtendAI approach, into Maritime Rescue Coordination Centers. In this framework, the human operator first develops and explains an operational plan, after which the AI extends the reasoning by providing structured feedback, identifying cognitive gaps, and supporting reflection while keeping the final decision under human control. The model combines LLM interaction, optimization methods, and supporting datasets to enhance planning quality. Reported benefits include increased decision confidence and satisfaction, as well as improved reasoning depth, though at the cost of greater cognitive effort (the results are drawn from the main study Reicherts et al., not from primary data). Overall, the study argues that AI should function as a complementary reasoning partner, improving outcomes without replacing human judgment in critical MSAR operations.  

D. Papachristos, N. Nikitakos, A. Gegenava et al. · 0 citations
Preprint Jul 2026

End-to-End LLM Flight Planning with RAG-based Memory and Multi-modal Coach Agent

Bridging the gap between human pilot intent and autonomous flight operation is critical for real-world electric vertical takeoff and landing (eVTOL) aircraft deployment. Flight planning traditionally relies on classic algorithms that struggle to incorporate flexible human preferences. We present FRAMe, an End-to-End Large Language Model (LLM) Flight Planning tool with RAG-based Memory and Multi-modal Coach Agent. Our system integrates a planner LLM with a multi-modal coach agent and retrieval augmented generation (RAG)-based memory to generate flight plans that satisfy mission constraints while aligning with human flight operator preferences. We demonstrate the system in a range of real-world-inspired scenarios of varying difficulty levels. Across four LLMs, the full FRAMe system (RAG and coach) yields the highest validity for every planner (up to 93.8% aggregate, 99% on Easy scenarios for the strongest planner) and shifts preference-relevant metrics in the operator-favored direction where the metric has headroom. FRAMe signifies how advanced LLMs can be deployed for human-centric mission planning, translating natural language instructions into safe, efficient, and flexible flight routes. The code is available at: github.com/amin-tabrizian/FlightPlanningLLMs

Amin Tabrizian, Arsyi Aziz, Aarifah Ullah et al. · 1 citation
Book Open access Jul 2026

Context Aware AI Assistant and AR Interface for Lunar Extravehicular Activity (EVA) Procedural Guidance

As human space exploration returns to the Moon, astronauts need rapid access to procedural information during extravehicular activities (EVAs), where attention is divided across navigation, repair tasks, tool handling, and environmental risk. The challenge is not the absence of information, but surfacing the right information at the right moment. We present GAIN-AI (Guided Assistant for Intelligent Navigation), a context-aware AI assistant and minimal heads-up interface for procedural guidance in simulated lunar EVA. The system operates in two layers. The first grounds a large language model with structured context: EVA procedure documents, live telemetry data, and error-handling protocols encoded as JSON. The second restructures that output into three compact units for AR display: Goal, Task, and Verification. Evaluated on 111 synthetic EVA scenarios, the system scores 10.0/10 on nominal conditions and 8.15/10 on single-fault scenarios, with performance degrading on multi-fault and boundary-threshold cases.

Rodrigo Gallardo, Qilmeg Doudatcz, Ganit Goldstein et al. · 0 citations
Open access 2026

Automated Generation of Situational Judgment Tests for Civil Aviation Flight Attendants Using Large Language Models: Method and Preliminary Evaluation

This study aims to construct and validate a retrieval-augmented generation (RAG)-driven workflow for automatically generating SJT items and provides preliminary evidence for the feasibility of an automated development pathway for psychological assessment tools based on LLMs and RAG technology.

Yaqian Liu, Qida Hao, Jian Cheng et al. · 0 citations