TY - GEN
T1 - Large Language Models Are Unreliable for Cyber Threat Intelligence
AU - Mezzi, Emanuele
AU - Massacci, Fabio
AU - Tuma, Katja
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2025.
PY - 2025
Y1 - 2025
N2 - Several recent works have argued that Large Language Models (LLMs) can be used to tame the data deluge in the cybersecurity field, by improving the automation of Cyber Threat Intelligence (CTI) tasks. This work presents an evaluation methodology that other than allowing to test LLMs on CTI tasks when using zero-shot learning, few-shot learning, and fine-tuning, also allows to quantify their consistency and their confidence level. We run experiments with three state-of-the-art LLMs and a dataset of 350 threat intelligence reports and present new evidence of potential security risks in relying on LLMs for CTI. We show how LLMs cannot guarantee sufficient performance on real-size reports while also being inconsistent and overconfident. Few-shot learning and fine-tuning only partially improve the results, thus posing doubts about the possibility of using LLMs for CTI scenarios, where labelled datasets are lacking and where confidence is a fundamental factor.
AB - Several recent works have argued that Large Language Models (LLMs) can be used to tame the data deluge in the cybersecurity field, by improving the automation of Cyber Threat Intelligence (CTI) tasks. This work presents an evaluation methodology that other than allowing to test LLMs on CTI tasks when using zero-shot learning, few-shot learning, and fine-tuning, also allows to quantify their consistency and their confidence level. We run experiments with three state-of-the-art LLMs and a dataset of 350 threat intelligence reports and present new evidence of potential security risks in relying on LLMs for CTI. We show how LLMs cannot guarantee sufficient performance on real-size reports while also being inconsistent and overconfident. Few-shot learning and fine-tuning only partially improve the results, thus posing doubts about the possibility of using LLMs for CTI scenarios, where labelled datasets are lacking and where confidence is a fundamental factor.
UR - https://www.scopus.com/pages/publications/105014146093
UR - https://www.scopus.com/pages/publications/105014146093#tab=citedBy
U2 - 10.1007/978-3-032-00627-1_17
DO - 10.1007/978-3-032-00627-1_17
M3 - Conference contribution
AN - SCOPUS:105014146093
SN - 9783032006264
VL - 2
T3 - Lecture Notes in Computer Science
SP - 343
EP - 364
BT - Availability, Reliability and Security
A2 - Dalla Preda, Mila
A2 - Schrittwieser, Sebastian
A2 - Naessens, Vincent
A2 - De Sutter, Bjorn
PB - Springer Nature
T2 - 20th International Conference on Availability, Reliability and Security, ARES 2025
Y2 - 11 August 2025 through 14 August 2025
ER -