top of page

Princeton Journal of Interdisciplinary Research, Volume 1, Issue 3

— Bridging Horizons (March 2026) - ISSN 3069-8200

How Effective are Large Language Models at Detecting Security Vulnerabilities?

Author: Stuart D. Tioniwar

Affiliation: Cambridge Centre for International Research

Abstract: With the prevalence of software systems and digitized information, (cyber) security vulnerabilities have also become an important focus. Large Language Models (LLMs) have been seen as potential agents in detecting these vulnerabilities. However, the question of "how effective can LLMs be in this role?" arises. This paper contributes a detailed evaluation of the LLMs' capabilities to carry out vulnerability detection (VD) along with the engineering of OverallPrompt (OP) and OverallPrompt-Few Shot (OP-FS), as well as the FunctionOnlyDataset in an attempt to improve the LLMs' performance. This evaluation is measured by multiple metrics, including accuracy, F1, and Bias rate; focusing on quantitative binary results. Through meticulous analysis, this research paints a clear picture that LLMs are not suitable for real-world use as currently constructed: the best performances peak at 0.57 accuracy and 0.67 F1 score. This unreliability is not good enough for real-life situations where incorrect VD detections can result in catastrophic consequences. However, the improvement in results when met with different prompts shows great promise in the capabilities of LLMs to further improve and result in more accurate detections by altering certain moving parts of the LLM network. Using various prompting, pretraining, and fine-tuning methods in the future can greatly impact the VD capabilities of LLMs and bring them further towards real-world reliability.

Keywords: LLM (Large Language Model), vulnerability detection, prompts, performance

The Princeton Journal of Interdisciplinary Research (PJIR) · ISSN 3069-8200

bottom of page