JavaVulBench: A Java Vulnerability Benchmark with Realistic Splits, a Unified Multi-Backend Harness, and a Leakage-Aware Evaluation Mode 文章

ArXiv CS.AI2026-07-07PAPERen作者: Norbert Sandor Szolnoki, Gabor Antal

详细信息

来源站点
ArXiv CS.AI
作者
Norbert Sandor Szolnoki, Gabor Antal
文章类型
PAPER
语言
en
发布日期
2026-07-07

摘要

arXiv:2607.02825v1 Announce Type: cross Abstract: We release \textsc{JavaVulBench}, a benchmark dataset and evaluation harness for Java vulnerability detection. The dataset contains $\sim$30{,}600 Java methods spanning 1{,}740 CVEs and 700+ projects, labelled at both method and line granularity, with per-CVE publication dates and five realistic split strategies: random, project-disjoint, temporal, deduplicated, and unseen CWE-family. The harness provides a single \texttt{LlmPrediction} schema across three backend families (encoder classifiers, local generative models served by Ollama, and API-served LLMs routed through OpenRouter) so that twelve reference detectors CodeBERT, GraphCodeBERT, UniXcoder, DeepSeek-Coder-1.3B, and eight API/open-weight LLMs (GPT-4o, GPT-4.1-mini, Claude Sonnet~4, DeepSeek-v3, DeepSeek-Coder-v2, Qwen-2.5-Coder-14B/7B, CodeLlama-13B) are evaluated under identical conditions from a single command.

相关事件

暂无数据

相关公司查看全部 (2)

A
AMI团队RESEARCH_INSTITUTE
A
AT TCOMPANY

相关人物

暂无数据