Automated reproducible malware analysis: A standardized testbed for prompt-driven LLMs
2025 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE credits
Student thesis
Abstract [en]
Malware poses a persistent threat to every digitalized entity in cyberspace. As software becomes ubiquitous in essential functions and critical infrastructure, malicious code becomes a risk even to societal continuity. It is a threat that is constantly expanding and evolving in lockstep with other technological advancements. Malware analysis is a practice that seeks to derive valuable insights regarding the threat’s semantics and semiotics. Such intelligence is vital to countering the proliferation of malicious software. However, manual analysis is not scalable with the exponential growth of software-based utilities. Consequently, the demand for automation has become more prevalent, and Machine Learning (ML) solutions have demonstrated promising performance. Likewise, the development of Large Language Models (LLMs) and their advanced capabilities has recently garnered significant interest.
Nevertheless, many contemporary models are only accessible through their prompt, referred to as prompt-driven LLMs. Interactions with these models involve submitting and receiving messages in natural language. This characteristic renders the performance evaluation more intricate because their responses are neither naturally nor easily quantifiable. Nonetheless, this study presents a fully automated framework for prompt-driven LLM evaluation, including automatic response quantification. Due to the lack of available quality data in malware analysis research, the framework also includes automated dataset development. It provides the study with high-quality contemporary data aggregated using real-world malware and state-of-the-art threat intelligence. This dataset is then used with the framework to evaluate four lightweight LLMs’ performance in static malware analyses.
The result is a malware classification correctness of 97.6% with the average response time of 4.7 seconds and negligible 0.4‰ invalid responses. Additionally, framework conformity regarding three out of four evaluated models is on average 99.6%. Thus, solidifying the presented framework as a valid evaluation tool for prompt-driven LLMs. These results provide valuable insight into the possibilities of LLMs in cybersecurity. Meanwhile, it provides an extensive automated framework enabling future threat intelligence and protection research. Furthermore, this study presents a novel approach to performance quantification of prompt-driven LLMs. Indeed, a fully automated framework for static malware analysis and a proposal for a standardized testbed.
Place, publisher, year, edition, pages
2025. , p. 55
Keywords [en]
Large language models (LLMs), malware analysis, reproducibility, standardized testbed, automated framework, lightweight LLMs, open-source, cybersecurity
National Category
Information Systems, Social aspects
Identifiers
URN: urn:nbn:se:his:diva-25315OAI: oai:DiVA.org:his-25315DiVA, id: diva2:1973561
Subject / course
Informationsteknologi
Educational program
Privacy, Information and Cyber Security - Master's Programme 120 ECTS
Supervisors
Examiners
2025-06-192025-06-192025-09-29Bibliographically approved