Consensus Generation of Multi-Vehicle Accident Reports Based on Vision-Language Models and Graph Neural Networks
As road traffic accidents are high-frequency incidents. Although there are sufficient and clear traffic control cameras at some perception-dense sites, in other places there is often a lack of enough cameras. In order to make full use of limited information to efficiently obtain an overview of the accident scene in a short time, assist on-site handling and investigation, and help reduce secondary casualties and restore traffic. This paper makes full use of existing technologies and builds a framework that can fully utilize the image and text information of the accident scene. First, the images uploaded by vehicles are used for VLM to generate objective descriptions, and the descriptions from multiple sensors within a single vehicle are fused. Then the network architecture is designed to carry out credibility calculation and severity classification. Finally, the summary is selected to form the final accident consensus report. Through the above work, semantic description and severity assessment of the event can be quickly achieved, while false reports are minimized to the greatest extent.