Improving Security Bug Report Classification with Large Language Models

Published on
Embed video
Share video
Ask about this video

Scene 1 (0s)

[Audio] Libyan Academy – Jabal Al Akhdar Branch School of Basic Sciences | Department of Computer Science Improving Security Bug Report Classification with Large Language Models Master's Thesis Progress Seminar Presented by Riad Dawood Supervised by Dr Samia Abdalhamid.

Scene 2 (21s)

[Audio] Introduction Software development follows the Software Development Life Cycle (S-D-L-C-), where Maintenance focuses on fixing bugs and addressing security vulnerabilities. Bug Reports document software defects, but their unstructured text makes automatic classification challenging. Traditional classifiers often fail to generalize across software projects due to limited contextual understanding. Recent L-L-M's offer powerful contextual understanding, making them a promising solution for S-B-R classification. Motivation The effectiveness of L-L-M's and different prompting strategies for cross project S-B-R classification remains insufficiently explored..

Scene 3 (1m 8s)

Feature-Based Classification • N-Gram IDF (Terdchrisnari et al., 2017) [8] • Label Quality Impact (Afric et al., 2023) [10] • Cross-Project Learning (Sharma et al., 2012) [17].

Scene 4 (1m 42s)

Gap I Limited Context Understanding. Research Gap.

Scene 5 (1m 50s)

Weak Context Understanding. Problem Statement. ProjectA X Project B Limited Generalization.

Scene 6 (1m 59s)

01 Evaluating LLM-based approaches for Security Bug Report (SBR) classification and their effectiveness compared with traditional machine learning methods. Traditional ML Methods LLM-based Approaches Metrics: Accuracy, Precision, Recall, Fl-score.

Scene 7 (2m 20s)

How effective are large language models in classifying security bug reports? RQI: Collect Security Bug Reports LLMs Receive Bug Reports for Classification Efficiency Evaluation GitLab Llama 3 ClauO 3 Evaluate Overan efficiency of LLMs in classifying security bug reports. Gemini.

Scene 8 (2m 51s)

PHASE IV Cornparison with CNN Baseline CNN Baseline CNN Modet (I D Convolution + Dense) Sarne Preprocessed Datasets Performance Evaluation M etrics Precision Recall F I — score Trained using Stratified 10 —fold cross—validation Compare Results (LLMs vs CNN).

Scene 9 (3m 6s)

[Audio] Dataset Description Project Reports SBRs % S-B-R's Ambari 1000 56 5.61% Camel 1000 74 7.40% Derby 1000179 17.9% Wicket 1000 47 4.70% Chromium 41940807 1.96% Total 45940 1163 2.53% Total Security Bug Reports (SBRs) 1163 Total Reports 45940.

Scene 10 (4m 1s)

[Audio] Models Evaluated Large Language Models (LLMs) Mistral-7B (Local) 7B Parameters DeepSeek chat (A-P-I--) 671B Mixture of Experts (37B Active) GPT-3.5 turbo (A-P-I--) Closed source (Not disclosed).

Scene 11 (4m 22s)

undefined. Models Evaluated Baseline Model CNN (Traditional Baseline).

Scene 12 (4m 41s)

[Audio] Prompting Strategies ToT Tree of Thoughts Explore multiple reasoning paths before deciding. FS Few Shot 2-5 examples provided for better context. OS One Shot One example provided to guide The model. CoT Chain of Thought Step by step reasoning for better accuracy. ZS Zero Shot Direct instruction without examples..

Scene 13 (5m 5s)

[Audio] research results (Answers to R-Q-1-) DeepSeek Chat is the best performing L-L-M (Avg. F1 = 0.90). Chromium achieved the highest project specific F1-score (0.96). GPT-3.5-Turbo and Mistral-7B achieved the same average F1-score (0.83)..

Scene 14 (5m 36s)

[Audio] research results (Answers to R-Q-2-) Zero Shot and Few Shot are the best performing prompting strategies (Avg. F1 = 0.85). Chain of Thought and Tree of Thoughts achieved moderate average F1-scores (0.80 and 0.77). One Shot is the lowest performing prompting strategy (Avg. F1 = 0.72)..

Scene 15 (6m 8s)

[Audio] research results (Answers to R-Q-3-) CNN (Baseline) is the best performing model (Avg. F1 = 0.91). DeepSeek Chat achieved a competitive average F1-score (0.90). DeepSeek Chat performed comparably to the C-N-N baseline with only a 0.01 F1 difference..

Scene 16 (6m 33s)

[Audio] Contributions Comprehensive Evaluation of Three L-L-M's for Security Bug Report Classification. Comparative analysis of five prompting strategies (ZS, OS, FS, CoT, ToT). Benchmark of L-L-M's Comparative to a C-N-N baseline. Identification of the best LLM–prompt combinations. Practical recommendations for applying L-L-M's in S-B-R classification..

Scene 17 (7m 1s)

[Audio] Published Research Paper Title: Security Bug Report Classification: A Comparative Study of Convolutional Neural Networks and Large Language Models Status: Accepted and Presented at MI-STA 2026 Conference: 2026 I-E-E-E 5th International Maghreb Meeting of the Conference on Sciences and Techniques of Automatic Control and Computer Engineering (MI-STA) April 2026 DOI:10.1109/MI-STA68962.2026.11511275.

Scene 18 (7m 50s)

Weeks 1 2 3 Task Name 1 Thesis Revision Proofreading Final Submission Timeline: 4 Weeks 2 3 4.

Scene 19 (7m 56s)

[Audio] Thank You Questions and Discussion Contact: riadalelwani@gmail.com Supervisor Assoc. Prof. Samia Wanees Abdalhamid.