Shock horror — AI-generated security patches fall short of actually solving all the problems they were meant to fix
Researchers tested AI-generated patches on six CVEs with poor success rates Many fixes failed, altered behavior, or introduced new vulnerabilities Guidance improved outcomes, leading to FLAWED evaluation harness release When using Generative Artificial Intelligence (GenAI) to fix vulnerabilities, security professionals are most of the time just robbing Peter to pay Paul, experts have warned. Researchers from 1Passwords Off-by-1 Labs analyzed fixes proposed by two frontier models - ChatGPT 5.5 at “medium” effort, and Claude Opus 4.8 at “high” effort. As an experiment, the researchers took six recently disclosed CVEs and produced 6,080 patches using two frontier, cyber-capable reasoning models. The results were underwhelming to say the least - of all the proposed patches, just a quarter (26%) fully resolved the issue. FLAWED work? This obviously leaves plenty to be desired, as half (49.3%) of the p...