Loading…
Thursday November 5, 2026 11:30am - 12:15pm PST
There has been considerable discussion on how to use AI to find vulnerabilities, but very little discussion on how to use it to classify vulnerabilities. Given the huge backlog of vulnerabilities in our systems, and the impending agentic coding revolution which will 100x them, a new approach is needed to accurately cull and rank issues. In this talk, we discuss agentic classification vs. supervised learning-based classification, what other traits can be discerned besides simple "true or false positive", utilizing dynamic analysis techniques, frameworks for evaluation, confidence of results, strengths and weaknesses of generative AI in this task domain, and future research directions.

Problem Statement:

Security tools — SAST, DAST, IAST, SCA, and others — produce findings. Humans classify them. That process does not scale, and in practice it mostly doesn't happen — the majority of findings across the industry are never reviewed. False positive rates vary wildly by tool, rule, and codebase, but the deeper issue is that even *true* positives require judgment: is this exploitable in context? Are there compensating controls elsewhere? Is the reported severity accurate? This classification step — not detection — is where application security breaks down, and the problem compounds as AI-assisted code generation increases finding volume.

We set out to answer a practical question: what does it actually take to classify vulnerability findings with the accuracy and nuance of a senior security engineer? To find out, we ran a systematic bakeoff across a range of approaches — naive LLM prompting, supervised learning, multiple agentic architectures with different reasoning strategies, domain knowledge bases, and dynamic analysis techniques — evaluated against benchmarks constructed from real findings in real organizations across 15+ security tools. We compared structured decision trees against open-ended ReACT reasoning, tested how much domain-specific knowledge bases improve accuracy, measured the limits of attention-based analysis on complex multi-file data flows, and assessed when dynamic exploit verification is worth its cost. The results show where each approach succeeds, where it fails, and what combination gets closest to expert-level classification.
Speakers
avatar for Arshan Dabirsiaghi

Arshan Dabirsiaghi

CTO and Co-Founder, Pixee
Arshan is a security researcher and developer pretending to be a software executive, with many years of experience advising large organizations on code security and building tooling to support secure code development. He has spoken at prestigious conferences like Blackhat and OWASP... Read More →
avatar for Ryan Dens

Ryan Dens

Software Engineer, Pixee
Ryan is a software engineer passionate about security and developer productivity
   linkedin.com/in/ryan-dens/
 ryandens.com (blog)
 pixee.ai (company)
... Read More →
Thursday November 5, 2026 11:30am - 12:15pm PST
Room: Seacliff AB (Bay Level)

Sign up or log in to save this to your schedule, view media, leave feedback and see who's attending!

Share Modal

Share this link via

Or copy link