Skip to content

40+ AI Tool Security Vulnerability Statistics (2026 Data)

Matt Li By Matt Li Co-Founder and Director 13 min read
TL;DR: AI tools now show up in breach reports, CVE feeds and the US government's list of exploited flaws. In IBM's 2026 Cost of a Data Breach study, 21% of organizations had a breach involving an AI model or application, and 92% of those lacked proper AI access controls. Veracode's tests found AI models write code that passes security checks 56% of the time.

Langflow, an open-source tool for building AI agents, has six flaws in CISA’s catalog of vulnerabilities attackers are exploiting in the wild.

The first, CVE-2025-3248, let anyone who could reach the server run code on it without logging in. CISA lists it as used in ransomware campaigns and added the other five between March and August 2026.

Key takeaways
  1. 1By our NVD query, 75 CVEs published in 2026 through September 13 mention prompt injection in their descriptions, against 34 in all of 2025.
  2. 2OWASP’s 2026 LLM Top 10 moved Excessive Agency from sixth to third, the biggest climb on the list.
  3. 3Security incidents involving shadow AI more than doubled in IBM’s data, to 43% from 20% a year earlier.
  4. 4Gartner expects spending on securing AI to reach $4.8 billion in 2027, up 68.7% on 2026.

Which AI Tools Have Flaws Attackers Are Exploiting?

Eleven flaws in AI development tools entered the CISA Known Exploited Vulnerabilities catalog between January 1 and September 13, 2026. Only one did in 2025, and none of the 12 sits in a chatbot.

They sit in the platforms engineers use to build agents, route model traffic and run training jobs.

Added to KEVToolCVEWhat an attacker getsCVSS 3.x
May 5, 2025LangflowCVE-2025-3248Code execution with no login9.8
Mar 11, 2026n8nCVE-2025-68613Code execution through workflow expressions8.8
Mar 25, 2026LangflowCVE-2026-33017Code injection through public flows, no login9.8
May 8, 2026LiteLLMCVE-2026-42208SQL injection into the proxy database9.8
May 21, 2026LangflowCVE-2025-34291Account takeover, then code execution8.8
Jun 8, 2026LiteLLMCVE-2026-42271Command execution by low-privilege key holders8.8
Jul 7, 2026LangflowCVE-2026-55255Running flows that belong to other users8.4
Jul 21, 2026LangflowCVE-2026-0770Remote code execution9.8
Aug 4, 2026LangflowCVE-2026-9198Superuser token, then code execution, no login9.8
Aug 17, 2026RayCVE-2025-62593Code execution through a developer’s browser8.8
Aug 19, 2026MLflowCVE-2026-64849Requests to internal and cloud metadata services9.3
Sep 2, 2026LiteLLMCVE-2026-59822Access to the MCP endpoint with a fake auth header8.2
How we counted. We filtered the KEV feed of September 11, 2026 (1,709 entries) for AI agent, model gateway and machine learning platforms. Scores are NVD’s CVSS 3.x where NVD has scored the flaw, otherwise the score of the organization that filed it.

Langflow alone accounts for six entries and LiteLLM, a proxy that routes calls to model APIs, for three. By our reading of the NVD descriptions, 8 of the 12 end in code or command execution.

LiteLLM’s September entry is the first KEV listing tied to an MCP endpoint, the protocol AI agents use to call tools.

Patching has not kept up. Verizon’s 2026 Data Breach Investigations Report found organizations fully fixed only 26% of KEV-listed flaws in 2025, down from 38%. The median time to full resolution rose to 43 days from 32.

Exploiting a vulnerability is now the most common initial access vector in breaches, at 31%.

How Many CVEs Hit AI Coding Assistants and MCP Servers?

We queried the NVD API for CVE records whose descriptions contain the phrase “prompt injection”. The count more than doubled in 2026 before the year was three-quarters done.

Records naming the Model Context Protocol went from 23 in 2025 to 53 by September 13, 2026.

Column chart of CVE records whose NVD description contains the phrase prompt injection, by publication year: 2 in 2023, 18 in 2024, 34 in 2025 and 75 in 2026 through September 13, from an NVD API query on September 14, 2026.

A keyword count misses bugs described in other words, so treat it as a floor. The cards show six tools developers run on their own machines, each with a flaw scored high or critical.

CursorAug 2025
9.8
CVE-2025-54135: the agent could create a new MCP config file with no approval. Fixed in 1.3.9.
Claude CodeAug 2025
9.8
CVE-2025-54795: a command parsing error skipped the confirmation prompt. Fixed in 1.0.20.
mcp-remoteJul 2025
9.6
CVE-2025-6514: OS command injection when connecting to an untrusted MCP server.
MCP InspectorJun 2025
9.4
CVE-2025-49596: no authentication between client and proxy. Fixed in 0.14.1. CVSS 4.0.
mcp-server-gitDec 2025
9.1
CVE-2025-68145: the repository restriction flag was not enforced on later tool calls.
GitHub CopilotAug 2025
7.8
CVE-2025-53773: command injection in Copilot and Visual Studio lets an attacker run code locally.
Scores and descriptions: NVD records, read September 14, 2026. NVD’s score shown where it has one; the filer’s score otherwise.

The first zero-click case against a production assistant was EchoLeak, logged as CVE-2025-32711 in June 2025. A crafted email could make Microsoft 365 Copilot send internal data to an attacker without the user clicking anything.

Our GitHub Copilot statistics and Cursor statistics show how many developers run these tools.

Severity depends on who scores it

EchoLeak carries two CVSS 3.1 scores on its NVD record: 9.3 from Microsoft and 7.5 from NVD’s analysts. For Cursor’s MCP flaw the gap runs the other way, 8.5 from the filer and 9.8 from NVD.

A patch queue sorted on one score can put the same bug in a different week.

Dot plot comparing CVSS 3.1 base scores from the filing vendor or CNA with NVD scores for the same flaw: Cursor CVE-2025-54135 8.5 vs 9.8, Cursor CVE-2025-54136 7.2 vs 8.8, Microsoft 365 Copilot CVE-2025-32711 9.3 vs 7.5, and n8n CVE-2025-68613 9.9 vs 8.8.

What Are the Top Security Risks for LLM Applications?

Prompt injection is still first on the OWASP Top 10 for LLM Applications 2026, published in August. Ranked by public incident data alone, it would drop out of the top 10.

The project leads keep it first because teams fight injection hard, so fewer clean exploits reach public databases.

Practitioner vote
  • Carries 75% of the weight in the final ranking
  • Ranks prompt injection #1 and misinformation near the bottom
Incident record
  • 7,714 incidents collected, 6,639 classified, 25% of the weight
  • Puts prompt injection outside the top 10 and misinformation near the top
Source: OWASP Top 10 for LLM Applications 2026, letter from the project leads.

The order moved more than in past editions. Excessive Agency, where a model has more tools or permissions than its task needs, climbed three places because agent deployments are where damage is landing.

Unbounded Consumption rose four places. Improper Output Handling fell the furthest, from fifth to tenth.

Slope chart of OWASP Top 10 for LLM Applications ranks in 2025 and 2026: prompt injection 1 to 1, sensitive information disclosure 2 to 2, supply chain 3 to 4, data and model poisoning 4 to 5, improper output handling 5 to 10, excessive agency 6 to 3, system prompt leakage 7 renamed hidden context exposure at 8, vector and embedding weaknesses 8 to 9, misinformation 9 to 7, unbounded consumption 10 to 6.

System Prompt Leakage is gone as a name. Its replacement, Hidden Context Exposure, also covers tool schemas, retrieved policy text and developer instructions.

Several entries absorbed newer attacks: prompt injection now includes instructions hidden in images or audio, and output handling now covers the insecure code assistants generate at scale.

“Stop trying to build a model that cannot be fooled. Build the system around it, so that when the model is fooled, and it will be, nothing important breaks.”

Steve Wilson and Rock Lambros, OWASP Top 10 for LLM Applications 2026

Once a model can call tools and act on its own, OWASP points readers to a separate list. The Top 10 for Agentic Applications came out on December 9, 2025, built with more than 100 contributors.

Our AI agents statistics track how fast those deployments are spreading.

How Much Do AI-Related Breaches Cost?

A breach involving model inversion, where an attacker pulls sensitive data back out of a model, cost $6.07 million on average in the IBM Cost of a Data Breach Report 2026. That is 18% above the global average.

Ponemon Institute studied 602 breached organizations for the report, covering incidents from March 2025 to February 2026.

Bar chart of average breach cost by AI incident type in the IBM Cost of a Data Breach Report 2026, in US dollars: model inversion 6.07 million, prompt injection 5.89 million, cloud misconfiguration of AI workloads 5.25 million, global average for all breaches 4.99 million, malicious model 4.94 million and model evasion 4.72 million.

The two most common entry points sat around the model, not in it: compromised APIs, apps or plug-ins (27%) and cloud misconfigurations affecting AI workloads (27%).

Breaches of open-source models averaged $5.63 million, against $4.98 million for models trained in-house.

40%
of organizations use access controls on AI models and data
68%
of breached organizations lacked AI governance, up from 63%
19%
coordinate AI governance with their security teams
$5.39M
average cost of a shadow AI incident, up from $4.63M
Source: IBM Cost of a Data Breach Report 2026, July 2026.

The share of organizations with an AI-related breach rose from 13% to 21%, a 61% increase by IBM’s own calculation. A year earlier, IBM’s 2025 report put that figure at 97% of organizations with an AI breach.

About one in five shadow AI incidents in 2026 ended in a regulatory fine.

Defenders use AI too. IBM found that organizations using AI and automation in security operations cut breach costs by almost $2 million on average.

SentinelOne’s guide to what artificial intelligence in cybersecurity can and cannot do weighs those gains against the risks.

Is AI-Generated Code Secure?

Across more than 100 models tested over four years, Veracode’s security pass rate has not moved: 55% in its first report and 56% in the 2026 GenAI Code Security Report. The 2026 round added 11 new models across 80 tasks.

Models now write code with correct syntax close to 100% of the time. Even so, about 44% of the coding tasks in Veracode’s tests produced a known vulnerability.

Four donut charts of the average security pass rate for AI-generated code by flaw type in Veracode's 2026 GenAI Code Security Report: cryptographic algorithms 87 percent, SQL injection 83 percent, cross-site scripting 15 percent and log injection 12 percent.

Models handle SQL injection and cryptography well and fail most cross-site scripting and log injection tasks. Java is the weakest language, with a 30% pass rate.

GPT-5.5 led the summer 2026 group at 68%, while six of the 11 new models scored between 50% and 53%.

  • Coding models: models built for code averaged 51%, general-purpose models 52%.
  • Reasoning: reasoning models averaged 56%, non-reasoning models 51%.
  • Size: large models averaged 53%, medium and small models 51% each.

Secrets are the other leak. GitGuardian counted 28.65 million new hardcoded secrets in public GitHub commits in 2025. Keys for AI services rose 81% to 1,275,105.

Commits made with Claude Code leaked secrets at a 3.2% rate, against 1.5% across all public commits.

MCP configuration files held 24,008 unique secrets on public GitHub, 2,117 of them valid credentials. For defect rates beyond security, see our AI-generated code quality statistics and the wider AI in software development statistics.

How Do Employees Leak Data Into AI Tools?

Source code is the data employees most often paste into outside AI models, by a large margin.

Verizon’s 2026 DBIR found shadow AI became the third most common non-malicious insider action in its data loss prevention dataset in 2025, four times its share a year earlier.

45% of employees use AI on work devices, up from 15%67% of AI users sign in with non-corporate accounts15%+ of users have unauthorized AI browser extensions3.2% of AI DLP violations involve research documents

Verizon counts a regular user as someone who opens an AI platform at least once every 15 days. The share signing in with personal accounts fell to 67% from 72%.

Browser extensions are a newer route out: one main job of these AI plug-ins is collecting what the user browses, internal sites included.

The best-known leak predates all of this. In April 2023 Samsung staff leaked internal data into ChatGPT, and Samsung banned generative AI tools on company devices from May 1.

In its own staff survey, about 65% said the tools carry a security risk. Our shadow AI statistics cover unapproved use in more detail.

How Are Attackers Using AI Tools?

In September 2025 a Chinese state-sponsored group used Claude Code to run an espionage campaign against about 30 targets.

Anthropic reported that AI performed 80% to 90% of the work, with humans stepping in at 4 to 6 decision points per campaign. A small number of intrusions succeeded.

Apr to May 2023
Samsung: internal data leaks into ChatGPT; generative AI banned on company devices from May 1.
May 5, 2025
Langflow CVE-2025-3248 enters the KEV catalog, the first of the 12 AI tool entries in the table above.
Jun 2025
EchoLeak: zero-click data theft through Microsoft 365 Copilot, CVE-2025-32711.
Jul 2025
Amazon Q Developer: malicious code ships in VS Code extension 1.84.0 but fails to run because of a syntax error.
Aug 26, 2025
Nx on npm: poisoned packages live for 4 hours turn local AI command line tools against their owners.
Sep to Nov 2025
GTG-1002: Claude Code used against about 30 targets; Anthropic discloses it on November 13.
Sources: TechCrunch; CISA; arXiv; AWS security bulletin AWS-2025-015; Nx; Anthropic.

The Nx attack shows how an assistant becomes the payload. While scanning for secrets, the malware tried to use installed AI tools such as Claude and Gemini.

Wiz counted over a thousand valid GitHub tokens leaked and over 5,500 private repositories made public. The Amazon Q case was luckier: no customer environments changed.

AI-assisted malware leans on well-known techniques. The median threat actor in Verizon’s data used AI help for 15 documented techniques, and under 2.5% of AI-assisted malware involved rare methods.

Phishing made up 44% of AI-assisted initial access vectors.

In the phishing simulations Verizon analyzed, the median click rate on voice and text message lures was 40% higher than on email.

IBM puts the cost side at one in four malicious breaches being AI-enabled, a 56% rise, averaging $6 million per breach. Deepfake impersonation drove 45% of those attacks.

What Do Regulators and Analysts Expect Next?

The EU pushed back its cybersecurity rules for high-risk AI. Regulation (EU) 2026/1744, published on July 24, 2026, moved obligations for Annex III systems to December 2, 2027 and for AI in products covered by Annex I to August 2, 2028.

They had been due on August 2, 2026 and August 2, 2027.

Those obligations include Article 15 of the AI Act, which lists data poisoning, model poisoning, adversarial examples and confidentiality attacks among the attacks security measures should address.

Providers of general-purpose models with systemic risk have owed “an adequate level of cybersecurity protection” since August 2, 2025, and the omnibus left that date unchanged.

Gartner’s forecasts point to more incidents. By 2028, it expects 25% of enterprise generative AI applications to have at least five minor security incidents a year, up from 9% in 2025.

It points to MCP’s design, which favors interoperability and developer speed over security enforcement by default.

“MCP was built for interoperability, ease of use and flexibility first, so security mistakes can manifest without continuous oversight for agentic AI.”

Aaron Lord, Sr. Director Analyst, Gartner, April 9, 2026

For major incidents, Gartner forecasts 15% of those applications a year by 2029, up from 3%. It also predicts that by 2029 over half of successful attacks on AI agents will exploit access control weaknesses and prompt injection.

Its survey of 316 executives ranked AI-driven discovery of vulnerabilities as the top emerging risk in the second quarter of 2026.

Budgets follow the risk. Gartner puts spending on securing AI at $2.8 billion in 2026 and nearly $7.7 billion by 2028. In its 2027 forecast, AI usage control grows fastest, at 73%.

Treemap of Gartner's 2027 forecast for spending on securing AI, 4,783 million US dollars in total: other securing AI 2,292 million, AI application security 851 million, AI usage control 749 million, AI governance platforms 462 million and AI gateways 429 million.

Application security stays the largest named segment in 2027 at $851 million. Gartner expects big security vendors to buy the startups crowding into application security and usage control.

Hiring Engineers Who Can Secure AI Systems

Most of the flaws above sit in the code and infrastructure around a model: gateways, agent builders, MCP servers and CI secrets.

Second Talent matches companies with pre-vetted AI developers, AI agent developers and DevOps engineers who build and lock down these systems.

Our cybersecurity engineer rate card for Vietnam shows what the role costs. Tell us what you are building and we will send matching profiles.

Frequently Asked Questions

What is prompt injection?

It is an attack that slips instructions into text, images or files a model reads, so the model follows the attacker instead of the user. It can arrive in the chat itself or hidden in an email, web page or document the assistant processes.

What is the CISA KEV catalog?

It is the US Cybersecurity and Infrastructure Security Agency’s list of vulnerabilities with evidence of active exploitation.

US federal civilian agencies must patch listed flaws by set deadlines, and many companies use it to decide what to fix first.

Why can one CVE have two different CVSS scores?

The organization that files a CVE, often the vendor, scores it first. NVD analysts may add their own score from the same description, and they can judge attack complexity or impact differently. Both scores stay on the record.

Are AI coding assistants safe to use at work?

They have shipped serious flaws, and fixes are out for the ones listed here. The risk depends on the version you run, which MCP servers you connect and whether the agent can run commands without approval.

Treat their code like any unreviewed pull request.

Hiring developers in Southeast Asia?

Get Cost Guide
Matt Li

Written by

Matt Li is a tech-driven entrepreneur with deep expertise in global talent strategy, digital experience optimization, e-commerce, and Web3 innovation. He is the Co-Founder of Second Talent, a US-based company that connects businesses with top-tier tech professionals worldwide. Since launching the company in 2024, Matt has led its growth by leveraging technology to streamline remote hiring and scale distributed teams. With a background spanning product, operations, and innovation, Matt brings a cross-disciplinary perspective to the evolving digital economy. His work sits at the intersection of global talent, emerging technology, and scalable digital transformation.

More posts by Matt Li →

Keep Reading

How would you like to talk?

WhatsApp us Prefer texting at your own pace? Just hit us up on WhatsApp. We promise no spam and a hassle-free experience.

Loading available times…