The Bugs That Never Existed: How AI-Generated CVEs Got Critical Scores
JFrog Security Research examined six SQLite advisories published in a newly created GitHub repo (programmervuln/cveadvisory-). NVD flagged them as critical, and CISA’s ADP agreed. A broader audit of all 55 advisories from the same account found 54 fabricated, likely generated by an AI model. One contained a real bug wrapped in unverified CVE metadata.
The most severe finding, CVE-2026-51302, was initially scored 10.0 Critical by Red Hat. It described a use-after-free in a function, exprComputeOperands(), that didn’t exist in SQLite 3.41.0, the version it was reported against. That function was only added in mid-2025, more than two years after that release. The function cited as the root cause, sqlite3ReleaseTempReg(), doesn’t free memory at all. It recycles register indices, which makes the described bug architecturally impossible. The score has since been downgraded to 7.6.
How JFrog verified the claims
JFrog Security researchers cloned the official SQLite repository and checked out the exact versions referenced in each advisory: 3.41.0, 3.51.2, and 3.51.3. They compared the reported vulnerability mechanics against the real source code, then compiled the official releases inside isolated Docker containers to rule out environmental interference. Each advisory’s proof-of-concept SQL statements ran against the compiled binaries under AddressSanitizer, a tool built to catch memory bugs. Alongside that, they audited the CPE metadata and advisory fields across NVD and GHSA to check the tracking data itself.
More advisories, same pattern
CVE-2026-51303 claimed that ExprListDelete() fails to clear back-references in parent structures, causing a use-after-free supposedly patched in version 3.51.3. JFrog found no such back-reference pointers in the relevant SQLite structures and no trace of the fix. A diff between 3.51.2 and 3.51.3 showed zero changes to the relevant file. The proof-of-concept wasn’t even valid SQL and failed at the parsing stage.
CVE-2026-51296 cited a bug at lines 3555 and 3575 of json.c. In the target version, that file is 2,706 lines long. Neither line exists.
CVE-2026-51304 described a vulnerability in a function called with one argument. The real function takes two, including a database context pointer. SQLite also clears the relevant pointer immediately after deletion, closing off the described attack path.
Why this keeps happening
MITRE’s CVE submission process runs on a public form with no identity verification. Anyone can file a report and propose a severity score. Analysts at NVD used to manually analyze and enrich every new CVE, adding severity scores and affected product data, until a surge in report volume in February 2024 overwhelmed that process. CISA and other authorized publishers have tried to fill the gap, but the pipeline is fragmented and backlogged. Nothing in the current system requires a proof-of-concept or a working reproduction before a CVE gets published.
JFrog’s audit found 54 fabricated advisories out of 55 from the same source. Recurring red flags showed up across them: no mention on the vendor’s own security page, no linked commit or pull request, empty or contradictory metadata, and references to code that doesn’t exist in the reported version.
The alternative: don’t trust the score
So whatever gets through the CVE pipeline without proof lands in your triage queue as is. If your triage process trusts an incoming severity score at face value, a fabricated 10.0 can burn engineering hours chasing a bug that was never there. The later downgrade doesn’t fix it: 7.6 still puts a nonexistent bug in the High bucket. With an AI agent handling triage, it gets worse: the agent will search for the vulnerable function, draft a patch, and recommend changes to code that doesn’t exist.
How Whitespots Platform handles it
This scenario is covered in Whitespots Platform: automation rescores vulnerabilities before they reach the security team. Whitespots doesn’t treat every score as ground truth. Findings go through rule-based auto-validation before anything gets assigned. Deduplication runs on purpose-built rules that return the same result on every run. Custom CVSS rules rescore severity against what’s actually reachable in the codebase, so the advisory’s score gets checked before anyone acts on it. None of this depends on a CI/CD pipeline. One webhook on the VCS group is the entire integration, with zero pipeline edits.
The difference comes down to order. The AI approach stacks automation directly on top of a team’s gaps in process, tooling, and prioritization. Whitespots builds a working process first, then automates on top of it. The table below compares both approaches feature by feature, from validation and deduplication to cost at scale.
What AI is good for, and what it isn’t
AI finds bugs fast. It works well for the first few scans, hours, or weeks. Once cost and consistency enter the picture, the economics change: false positives climb, token spend scales with every commit, and the bill depends on usage patterns nobody set out to predict.
Whitespots doesn’t reject AI either. It runs a self-hosted LLM for specific steps, like explaining a finding. The validation and prioritization chain runs on rules, and teams decide which steps get AI on top. The price stays the same either way.
Small teams are the ones most likely to run security on an LLM alone. Whitespots has a Startup license built for them, with a flat annual price and the same auto-validation, deduplication, custom CVSS rules and pipelineless integration.
Comparison table: AI can find some bugs. It can’t run a process.
Any team can get findings from an LLM without much effort. Getting accurate, repeatable results at a predictable cost takes a rule-based process.
| Criteria | AI approach | Whitespots |
|---|---|---|
| Validation | Confirms user-suggested findings without independent check | Rule-based auto-validation |
| Deduplication | No dedicated deduplication logic. Depends on per-run model output | Purpose-built deduplication rules |
| False positive rate | No mechanism to catch drift as the codebase changes | Controlled via custom CVSS and validation rules |
| CI/CD dependency | Requires a scan call on every commit or PR to stay continuous | No CI/CD dependency. One webhook on the VCS group, zero pipeline edits |
| Cost at scale | Example: 100 repos, 10 commits/day, $0.10/commit is ~$100/day, and it climbs with every commit | €60,000/year, flat. Price doesn’t move with scan volume |
| Cost for small teams | Billed per scan from day one. Spend grows with the team and commit volume | Startup license: €24,000/year, flat, for up to 50 developers and 100 assets |
| Token spend over time | Scales with codebase size and commit volume, plus model pricing changes | Fixed annual price. No scaling cost as usage grows |
| Consistency | Same code can return different verdicts across separate runs | Rule-based validation logic. Consistent verdict across runs |
| Predictability | Total cost depends on commit volume and model pricing | Transparent, fixed pricing set upfront |


