Deep Learning and evidential reasoning for software vulnerability analysis
Abstract
Security and privacy are increasingly becoming the primary risk for many of the world’s enterprises, businesses, and consumers, as cyber criminals continue to exploit security flaws and weaknesses in the systems and technology that underpin modern life. To meet this concern, we need new approaches to improving software quality, performance, and reliability at an industrial scale. However, conventional approaches to this, such as Static Application Security Testing, often have very high false positive rates. In addition, it is somewhat simplistic to say whether software has a known vulnerability or not. Attackers typically analyse code from a weakness perspective, chaining these weaknesses across multiple components and employing techniques that exploit poor safety measures more than outright vulnerabilities. Therefore, a method is required that can detect these weaknesses, known as risk signals, chain or correlate them, focusing on discovering the relationships between them. This study investigates the use of deep learning, statistical tests and artificial intelligence reasoning techniques to identify software weaknesses such as CWEs from source code. Secondly, the proposed methods detect risk signals in the metadata associated with software commits to repositories such as GitHub, and we further explore the combination of the source code with its metadata to identify risky commits. Thirdly, we investigate the use of AI reasoning techniques to chain, or correlate, these risk signals horizontally across multiple software commits to recognise software weakness chains. This thesis demonstrates a strong potential to apply deep learning, statistical tests, and artificial intelligence reasoning techniques in software weakness detection while improving prediction accuracy. We provide a foundation for future studies on reasoning software weaknesses. This also helps developers understand the corresponding attack strategies behind risk signals by building up a more comprehensive view of the attack scenarios, allowing them to make timely decisions and take appropriate actions.