RegulaRAG is presented, a Retrieval-Augmented Generation pipeline that couples SmartChunking, reference-aware enrichment of paragraphs and tables via graph traversal, with Smart Retrieve&Rerank over these enriched units, and maintains strong performance, remaining stable even as the number of regulatory sources grows.
Abstract
Generating regulation-compliant test scenarios is essential for validating safety-critical automotive systems, yet Large Language Models (LLMs) struggle to ground outputs in long, hierarchical standards. We present RegulaRAG, a Retrieval-Augmented Generation (RAG) pipeline that couples SmartChunking, reference-aware enrichment of paragraphs and tables via graph traversal, with Smart Retrieve&Rerank over these enriched units. To test our system, we evaluate on a manually curated dataset covering all scenarios in UN Regulation No. 152 (AEBS). Our study comprises: (i) a three-step progressive search that identifies near-optimal retrieval parameters without exhaustive grid search; (ii) head-to-head comparisons against five baseline RAG systems; and (iii) a robustness stress test that scales the source corpus with distractor content. Outputs are evaluated using a customized penalized scoring metric. Across all experiments, RegulaRAG achieves the highest average Meta-Score (82.99), outperforming the next-best system by 43% (NoRAG: 57.94), while operating at 14k-25k tokens per query versus up to 500k for graphcentric baselines. It maintains strong performance, remaining stable even as the number of regulatory sources grows, whereas competing RAG systems degrade sharply in both quality and robustness.
A systematic empirical study of multiple strategies for context enrichment and optimization in LLM‐based unit test generation, conducted on seven diverse projects (three open‐source and four proprietary industrial systems), encompassing 261 distinct methods establish this optimized context strategy as a cost‐effective solution for scalable, industrial‐grade automated test generation.
Javier Ferrer, Francisco Chicano· Expert systems· 0 citations
Competence claims for a language model in a safety-critical domain are credible when measured against a standard the domain already enforces. We evaluate an open-weight 31-billion-parameter multimodal model (Gemma 4 31B-IT) on the U.S. Nuclear Regulatory Commission Reactor Operator Generic Fundamentals Examination (GFE), scoring it paper by paper against the 80% criterion applied to every human candidate, with no rounding up. The evaluation set is a census of every GFE administered at the March sitting from 2015 to 2021, giving seven pressurized water reactor (PWR) and seven boiling water reactor (BWR) papers and 697 scored items. Eight configurations cross three model states, the base model, supervised fine-tuning (SFT) on distilled chain-of-thought rationales and retrieval-augmented fine-tuning (RAFT), with three retrieval conditions, none and BM25 retrieval over the Department of Energy Fundamentals Handbooks under fixed-size and structure-aware chunking. Out of the box it answers 51.94% correctly and passes no paper. SFT with fixed-size chunking retrieval passes 8 of 14, reaching 80.23% on PWR items and 79.77% pooled, with a Wilson interval spanning the threshold. The preferred chunking granularity reverses with training state, structure-aware before fine-tuning and fixed-size after, so chunking optimized against a base model cannot be inherited by its fine-tuned descendant. RAFT trails SFT by 2.2 to 2.3 percentage points overall, and the deficit holds in all four reactor-type and chunking strata. The pipeline runs on one workstation with no network access at run time, and the result approaches operator-level command of engineering fundamentals without reliably achieving it.
A reproducible, human-validated evaluation framework applied to 13 strategies—four architectural families crossed with four reasoning variants crossed with four reasoning variants—across three SLMs spanning 3B–14B parameters, plus targeted ablations.
Balaji Venktesh, Amsaprabhaa M, G. Sundaram· International Conference on...· 0 citations
Retrieval-Augmented Generation (RAG) grounds Large Language Models (LLMs) in domain-specific knowledge, yet systematic comparisons of RAG architectures in non-English regulatory domains remain limited. We present a comparative study of context-window extension and Question-to-question Inverted Index Matching (QuIM-RAG), evaluated on the German Core energy market data register (Marktstammdatenregister). Both approaches are implemented in a fully local, privacy-compliant setup using open-source LLMs and assessed through parameter optimization, architectural comparison, and expert validation. QuIM-RAG achieves slightly higher factual correctness (F1: 0.56 vs. 0.54) in automatic evaluation. Expert evaluation shows that 65% of responses are factually correct and only 3% are incorrect. Our results further demonstrate that automated LLM-based metrics systematically underestimate performance and provide less interpretable outcomes compared to domain experts. The study contributes a systematic comparison of RAG strategies for German regulatory contexts, introduces an evaluation framework tailored to accuracy-critical domains, and outlines guidelines for privacy-compliant deployment of RAG systems in high-stakes settings.
Camillo Dobrovsky, Leon Schönberg, Sebastian Mirschel et al.· 2026 6th International Confe...· 0 citations
Span-Level Uncertainty Estimation (SLUE) is formalized, a new task that targets the natural granularity for uncertainty: semantically coherent text spans, each conveying a single assessable unit of meaning.
Yimeng Zhang, Yingying Zhuang, Ziyi Wang et al.· arXiv.org· 0 citations
Experiments across three instruction-tuned models show that HiRoute achieves high safety rates across multiple safety benchmarks while preserving safe-response helpfulness, reducing over-refusal, and maintaining competitive performance on general-purpose tasks.
Fangzhou Chen, Shiji Zhao, Mengyan Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.