Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance
DGEval is introduced, the first benchmark for evaluating LLM knowledge of IMDG Amendment 42-24, and shows that LLMs may support compliance tasks, particularly structured DGL lookups with web search, but unreliability in operational areas and regulatory-text recall means human oversight and authoritative source verification remain necessary before deployment in any safety-critical context.