Large-scale neural architectures exhibit systematic failures in compositional generalization and formal verifiability despite remarkable pattern recognition capabilities. This paper introduces the Neural-Symbolic-Verification (NSV) Loop–a functional decomposition framework–and uses it to systematically survey neuro-symbolic integration as a principled pathway toward artificial general intelligence. The NSV Loop organizes hybrid architectures through four computational stages structuring perception, symbolic execution, verification, and feedback. We operationalize the Grounding-Instructibility-Alignment (G-I-A) framework for production assessment and demonstrate quantifiable advantages: perfect compositional accuracy on SCAN (100% vs 13.8% neural baseline, length split), sample efficiency gains exceeding 10
$$\times $$
×
on visual reasoning tasks, and formal verification achieving certification rates above 95% with sub-100ms latency in autonomous systems. Analysis documents critical bottlenecks–grounding complexity scaling exponentially with entity count, cross-domain transfer exhibiting near-zero retention, and adversarial robustness evaluation remaining absent. The NSV+G-I-A framework enables systematic comparison manifesting when hybrid integration justifies complexity: safety-critical applications requiring formal guarantees, data-scarce environments, and compositional reasoning tasks. We establish clear capability boundaries distinguishing reliable improvements from speculative claims while proposing testable research directions with explicit validation protocols.
Safayat Bin Hakim, Kanchon Gharami, H. Wang et al.· Progress in Artificial Intel...· 0 citations
Security teams and researchers choose knowledge-graph extraction tooling for threat reports on the strength of published triple-F1 scores, yet those scores depend on how predicted triples are matched to gold annotations. We could reimplement the stated matching rule for only five of twelve inspected systems. Re-scoring ten system outputs on shared documents under eight protocols reverses eleven of forty-five pairwise orderings; one fixed prediction set spans 0.16-0.70 F1. On GRID's external 378-item calibration set, no mechanical matcher (lexical, embedding, or entailment) agrees with multi-reviewer adjudication above 71%, whereas an LLM judge reaches 86%. To separate component effects from matcher rewards, we build CTIForge, whose deterministic validation layer can vary while extraction is held byte-identical. Across seven tested deployment configurations, validation raises precision for all four hosted backbones and lowers it for all three offline backbones. Because backbone, decoding, and backend-specific prompting covary, this is a descriptive split rather than an isolated serving effect. It coincides with a roughly 2.8-fold increase in actions explicitly disputing entity type, consistent with hand-written rules encoding the conventions of the extractor against which they were developed. We release the pipeline, protocol suite, and per-triple audit records.
Safayat Bin Hakim, H. Song· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.