AI UPDATE: OpenAI’s GPT-5 line and DeepSeek’s V4 Flash are posting eye-watering results on CyberGym, the benchmark that measures whether an AI agent can find, reproduce, and exploit real software vulnerabilities. The headline numbers sound like progress. Whether they should feel like progress is a different question.
The Numbers, Fast
CyberGym tests agents against 1,507 real vulnerabilities pulled from 188 software projects, real bugs, real codebases, not toy problems. Recent published results put the top score at 86.9%, with several frontier models clustered close behind in the mid-80s. DeepSeek’s newly upgraded V4-Flash alone posted a 76.7% score on the benchmark.
For context, when CyberGym launched, the best score on record was around 20%. That’s not incremental improvement. That’s a step-change in what a piece of software can do to other software, unsupervised.
And the capability isn’t limited to spotting known issues. Independent analysis of the leaderboard notes that top models now force crashes in vulnerable targets nearly 99% of the time, though actually reproducing and diagnosing the specific underlying vulnerability still trails crash rates by a meaningful margin. In other words: these systems are extremely good at breaking things, and getting better, fast, at understanding exactly what they broke and why.
Here’s Where It Gets Uncomfortable
Every write-up on this frames it the same way: dual-use. Same tool, defense or offense, take your pick. That framing is technically true and also doing a lot of quiet work to make people feel okay about it.
Because here’s the thing nobody says out loud: the defenders and the attackers are not using this technology under the same constraints. A security team has to patch responsibly, get sign-off, follow disclosure timelines, work within budgets and headcount. An attacker just needs the exploit to work once. When you hand both sides a tool that can autonomously discover a zero-day, the side with fewer rules to follow gets the bigger relative advantage. That’s not a neutral outcome. That’s an asymmetry, and it favors whoever is willing to move fastest and care least.
What This Actually Means for Work
Set aside the hacking angle for a second and look at what this really demonstrates. These benchmarks aren’t just about security anymore, they’re a proof of concept for autonomous agents operating inside genuinely complex, high-stakes systems with minimal human oversight and getting it right the majority of the time.
That should make anyone doing complex knowledge work pay attention, not just security researchers. If a model can navigate a 188-project codebase it’s never seen, form a hypothesis about a hidden flaw, write a working proof of concept, and validate it, that is a demonstration of reasoning and autonomous execution under ambiguity. Vulnerability research was supposed to be one of the hardest, most human-intuition-dependent disciplines in software. It just got benchmarked at 80-90% success by machines. Whatever job you think requires deep contextual judgment and can’t be touched by this wave, it’s worth asking honestly what evidence you’re actually standing on.
The Trust Problem Nobody’s Pricing In
There’s also a quieter issue: verification. As these self-reported benchmark numbers proliferate across labs, several analysts have already flagged inconsistencies between vendor-published scores and independently reproduced ones, different harnesses, different conditions, numbers that shift depending on who’s measuring. We’re being asked to trust that a system is safe to deploy at scale based largely on the word of the company that built it and profits from the hype. That should give everyone pause, not just competitors trying to one-up each other on a leaderboard.
The Bottom Line
Better vulnerability detection genuinely could harden the software everyone depends on, banks, hospitals, power grids, all of it. That’s the optimistic case, and it’s real. But treating a 4x-to-5x jump in autonomous exploit capability as pure good news, without reckoning with who gets to use it first, how it changes the leverage between attackers and defenders, and what it signals about which human skills are next on the chopping block, is naive at best.
The technology isn’t good or bad. But pretending the humans deploying it are all equally careful, equally accountable, and equally slow-moving is the part of this story that deserves way more scrutiny than it’s getting.
Stay skeptical. Stay informed. Cointiculate.


