Virginia Tech® home

OmniVul: A Holistic, Multi-Turn Conversational Benchmark for LLM-Based Vulnerability Assessment

Research Paper Showcase 2026

Abstract

With more than 20,000 Common Vulnerabilities and Exposures (CVEs) reported annually, software vulnerabilities represent a critical cybersecurity challenge. This volume has intensified the demand for automated detection and analysis, motivating the integration of large language models (LLMs) for such tasks. However, existing vulnerability benchmarks are not suitable for evaluating LLMs' capabilities in vulnerability assessment, as most of them 1) rely on narrow data sources, 2) lack deep context, and 3) focus on single-turn Q&A rather than realistic, multi-stage analyst workflows. To address this gap, we introduce OmniVul, a comprehensive multi-turn benchmark for LLM-based vulnerability assessment. OmniVul comprises 2,000 CVEs with question–answer pairs spanning 23 attributes, including detection, code localization, root cause analysis, and patch suggestion. We employ an automated workflow to aggregate multi-source data via Retrieval-Augmented Generation (RAG), ensuring quality through LLM-as-a-Judge filtering and conformal prediction calibrated by human expert annotations. An evaluation of five state-of-the-art LLMs on OmniVul reveals distinct performance gaps, with top-1 accuracy remaining below 50% on average for vulnerable code detection and CVE identification. Our evaluation also demonstrates that current models lack critical reasoning capabilities for reliable vulnerability assessment. These results highlight the importance of OmniVul for advancing research in evaluating and fine-tuning LLMs for vulnerability assessment.


Authors

  • Vishnu Teja Kandalam, Ph.D. student, computer science, George Mason University
  • Viet Duong, Ph.D. student, computer science, College of William & Mary
  • Xiaochang Li, Ph.D. student, computer science, College of William & Mary
  • Minghui Yin, M.S. student, computer science, Cornell Tech
  • Vamsi Shankar Simhadri, Ph.D. student, computer science, George Mason University
  • Hung Pham, undergrad student, computer science, George Mason University
  • Huajie Shao, assistant professor, computer science, College of William & Mary
  • Xiaokuan Zhang, assistant professor, computer science, George Mason University
  • Yue Xiao, assistant professor, computer science, College of William & Mary

Publication

  • Venue: 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD '26)
  • Date: Aug 9, 2026

Related Papers

BOCLOAK: Optimal Transport-Guided Adversarial Attacks on Graph Neural Network-Based Bot Detection

A Comprehensive Assessment Tool for Prompt Injection Attacks in Large Language Models (LLMs)

An Empirical Study of LLM Serving in Confidential GPUs

FedHusky: Accelerating Hybrid Federated Learning with Client Hopping

MoltGraph: A Longitudinal Temporal Graph Dataset of Moltbook for Coordinated-Agent Detection

StealthInk: A Multi-bit and Stealthy Watermark for Large Language Models

Towards AI-Driven Human-Machine Co-Teaming for Adaptive and Agile Cyber Security Operation Centers