Wednesday, July 29, 2026
HomeSoftware DevelopmentVeracode Finds AI-Generated Code Safety Has Barely Improved Since Final Yr

Veracode Finds AI-Generated Code Safety Has Barely Improved Since Final Yr

-


Veracode’s 2026 GenAI Code Safety Report finds AI-generated code safety has stalled at a 56 % go price — with coding-specific fashions no safer than general-purpose ones.

  • AI fashions constructed for software program coding are not any safer than general-purpose fashions
  • Regardless of main good points in AI velocity, functionality, and reasoning, the report reveals almost half of AI-generated code nonetheless fails safety checks

Veracode, the worldwide chief in software threat administration, at present launched its 2026 GenAI Code Safety Report which reveals that, regardless of fast advances in AI coding capabilities, safety has stalled. Throughout 4 testing snapshots and greater than 100 fashions tracked for the reason that program started, the typical safety go price sits at 56 % — nearly unchanged since final 12 months’s report. Every mannequin was evaluated on code-generation duties spanning a number of programming languages and vulnerability classes, below standardized circumstances with no security-specific prompting. AI now generates roughly half of all dedicated code, however the safety hole isn’t narrowing.

The distinction is stark: fashions generate compilable code at a near-universal syntax go price of ~100%. In the case of safety, they fail almost 44 % of the time when given no security-specific steerage.

“As AI-fueled code velocity will increase, builders have gotten inundated with compliance dangers, safety alerts, and high quality points,” mentioned Chris Wysopal, Co-founder and Chief Safety Evangelist at Veracode. “We’re seeing a fast enhance within the adoption of AI-powered instruments to put in writing code and construct software program. However the root drawback stays: fashions could also be nearly syntactically excellent, however they’re nonetheless failing on almost half of all duties the place safety is required. That quantity needs to be a purple flag for any group.”

The GenAI Code Safety Report Leaderboard: Summer time 2026

Fig. 1: LLM Leaderboard by Security Pass RateFig. 1: LLM Leaderboard by Security Pass Rate

Fig. 1: LLM Leaderboard by Safety Move Fee

OpenAI’s GPT-5.5 leads this 12 months’s report at 68 %, whereas six of the 11 fashions cluster between 50 % and 53 %. Alibaba’s Qwen3.7-max comes final at 50 %, producing weak code each different output on this take a look at. The perfect mannequin obtainable at present nonetheless fails almost one in three safety duties. Notably, earlier editions of the report had been dominated by Western AI fashions. That’s not the case, with Kimi-K2.6 and MiMo-V2.5 outperforming a number of Western fashions. For enterprise safety groups, mannequin provenance is yet one more issue price weighing in for procurement choices, alongside safety go price.

Fashions Constructed for Coding and Bigger LLMs Aren’t Safer

Two broadly held assumptions don’t survive the info. First, fashions constructed particularly for writing software program code are not any safer than general-purpose AI. Coding-specialized fashions averaged a 51 % safety go price, in contrast with 52 % for general-purpose fashions. The checks had been run towards the uncooked fashions and never brokers or manufacturing environments with further tooling, guardrails, or human evaluation within the loop. This implies builders who go for coding-optimized instruments on the belief they’ll reduce safety threat are literally delivery weak code on the similar price as everybody else.

Fig. 2: Security Pass Rate of Coding-specific vs General-purpose LLMsFig. 2: Security Pass Rate of Coding-specific vs General-purpose LLMs

Fig. 2: Safety Move Fee of Coding-specific vs Common-purpose LLMs

Second, mannequin measurement has no affect on safety efficiency. Giant fashions (greater than 100 billion parameters) common 53 %, medium fashions common 51 %, and small fashions common 51 %. The one architectural issue that does assist: reasoning fashions preserve a constant safety edge over non-reasoning fashions, at 56 % versus 51 %. This means prolonged reasoning capabilities as a type of inside code evaluation.

Fig. 3: Security Pass Rate of Models by SizeFig. 3: Security Pass Rate of Models by Size

Fig. 3: Safety Move Fee of Fashions by Dimension

Java Is Nonetheless Final — However It’s Transferring within the Proper Path

Safety go charges by language vary from Python at 63 % all the way down to Java at solely 30 %. Java stays the riskiest language for AI code era by a big margin. Regardless of sub-optimal efficiency, it’s the solely language with a transparent, constant upward development over the previous 12 months.

“I’ve been vocal about making superior AI fashions, like Claude Fable and Mythos, obtainable to builders and defenders alike — and I stand by that place,” Wysopal closed. “The appropriate reply shouldn’t be limiting entry; it’s clear, evidence-based security. What this analysis makes clear is that AI-generated code must be handled like every unreviewed code: scan it, repair it, and by no means ship it blind. Till LLMs cause about safety the way in which they cause about syntax, guardrails within the improvement workflow aren’t elective.”

Managing Danger within the AI Period

Veracode recommends organizations take the next steps to handle the safety hole in AI-generated code:

  • Combine AI-powered instruments like Veracode Repair into developer workflows to remediate safety dangers in actual time.
  • Embed safety in agentic workflows to implement safe coding requirements robotically.
  • Use Software program Composition Evaluation (SCA) to detect vulnerabilities from third-party and open-source dependencies in AI-generated code.
  • Deploy Package deal Firewall to dam weak, malicious or non-compliant packages earlier than they attain the event setting.

To obtain the complete 2026 GenAI Code Safety Report, go to the Veracode web site. Attendees at Black Hat in Las Vegas August 4-7 are invited to go to sales space #4927 to study extra about AI code safety and the report’s findings.

About Veracode

Veracode is a world chief in Utility Danger Administration for the AI period. Powered by trillions of traces of code scans and a proprietary AI-assisted remediation engine, the Veracode platform is trusted by organizations worldwide to construct and preserve safe software program from code creation to cloud deployment. 1000’s of the world’s main improvement and safety groups use Veracode each second of every single day to get correct, actionable visibility of exploitable threat, obtain real-time vulnerability remediation, and cut back their safety debt at scale. Veracode is a multi-award-winning firm providing capabilities to safe all the software program improvement life cycle, together with Veracode Repair, Static Evaluation, Dynamic Evaluation, Software program Composition Evaluation, Container Safety, Utility Safety Posture Administration, Malicious Package deal Detection, Package deal Firewall, and Penetration Testing.

Be taught extra at www.veracode.com, on the Veracode weblog, and on LinkedIn and X.

Copyright © 2026 Veracode, Inc. All rights reserved. Veracode is a registered trademark of Veracode, Inc. in the USA and could also be registered in sure different jurisdictions. All different product names, manufacturers, or logos belong to their respective holders. All different emblems cited herein are property of their respective house owners.

Press and Media Contacts
Katy Gwilliam
Head of International Communications, Veracode
[email protected]

SD Times NewswireSD Times Newswire

Related articles

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Stay Connected

0FansLike
0FollowersFollow
0FollowersFollow
0SubscribersSubscribe

Latest posts