TL;DR: AI tools now show up in breach reports, CVE feeds and the US government's list of exploited flaws. In IBM's 2026 Cost of a Data Breach study, 21% of organizations had a breach involving an AI model or application, and 92% of those lacked proper AI access controls. Veracode's tests found AI models write code that passes security checks 56% of the time.
Langflow, an open-source tool for building AI agents, has six flaws in CISA’s catalog of vulnerabilities attackers are exploiting in the wild.
The first, CVE-2025-3248, let anyone who could reach the server run code on it without logging in. CISA lists it as used in ransomware campaigns and added the other five between March and August 2026.
- 1By our NVD query, 75 CVEs published in 2026 through September 13 mention prompt injection in their descriptions, against 34 in all of 2025.
- 2OWASP’s 2026 LLM Top 10 moved Excessive Agency from sixth to third, the biggest climb on the list.
- 3Security incidents involving shadow AI more than doubled in IBM’s data, to 43% from 20% a year earlier.
- 4Gartner expects spending on securing AI to reach $4.8 billion in 2027, up 68.7% on 2026.
Which AI Tools Have Flaws Attackers Are Exploiting?
Eleven flaws in AI development tools entered the CISA Known Exploited Vulnerabilities catalog between January 1 and September 13, 2026. Only one did in 2025, and none of the 12 sits in a chatbot.
They sit in the platforms engineers use to build agents, route model traffic and run training jobs.
| Added to KEV | Tool | CVE | What an attacker gets | CVSS 3.x |
|---|---|---|---|---|
| May 5, 2025 | Langflow | CVE-2025-3248 | Code execution with no login | 9.8 |
| Mar 11, 2026 | n8n | CVE-2025-68613 | Code execution through workflow expressions | 8.8 |
| Mar 25, 2026 | Langflow | CVE-2026-33017 | Code injection through public flows, no login | 9.8 |
| May 8, 2026 | LiteLLM | CVE-2026-42208 | SQL injection into the proxy database | 9.8 |
| May 21, 2026 | Langflow | CVE-2025-34291 | Account takeover, then code execution | 8.8 |
| Jun 8, 2026 | LiteLLM | CVE-2026-42271 | Command execution by low-privilege key holders | 8.8 |
| Jul 7, 2026 | Langflow | CVE-2026-55255 | Running flows that belong to other users | 8.4 |
| Jul 21, 2026 | Langflow | CVE-2026-0770 | Remote code execution | 9.8 |
| Aug 4, 2026 | Langflow | CVE-2026-9198 | Superuser token, then code execution, no login | 9.8 |
| Aug 17, 2026 | Ray | CVE-2025-62593 | Code execution through a developer’s browser | 8.8 |
| Aug 19, 2026 | MLflow | CVE-2026-64849 | Requests to internal and cloud metadata services | 9.3 |
| Sep 2, 2026 | LiteLLM | CVE-2026-59822 | Access to the MCP endpoint with a fake auth header | 8.2 |
Langflow alone accounts for six entries and LiteLLM, a proxy that routes calls to model APIs, for three. By our reading of the NVD descriptions, 8 of the 12 end in code or command execution.
LiteLLM’s September entry is the first KEV listing tied to an MCP endpoint, the protocol AI agents use to call tools.
Patching has not kept up. Verizon’s 2026 Data Breach Investigations Report found organizations fully fixed only 26% of KEV-listed flaws in 2025, down from 38%. The median time to full resolution rose to 43 days from 32.
Exploiting a vulnerability is now the most common initial access vector in breaches, at 31%.
How Many CVEs Hit AI Coding Assistants and MCP Servers?
We queried the NVD API for CVE records whose descriptions contain the phrase “prompt injection”. The count more than doubled in 2026 before the year was three-quarters done.
Records naming the Model Context Protocol went from 23 in 2025 to 53 by September 13, 2026.

A keyword count misses bugs described in other words, so treat it as a floor. The cards show six tools developers run on their own machines, each with a flaw scored high or critical.
The first zero-click case against a production assistant was EchoLeak, logged as CVE-2025-32711 in June 2025. A crafted email could make Microsoft 365 Copilot send internal data to an attacker without the user clicking anything.
Our GitHub Copilot statistics and Cursor statistics show how many developers run these tools.
Severity depends on who scores it
EchoLeak carries two CVSS 3.1 scores on its NVD record: 9.3 from Microsoft and 7.5 from NVD’s analysts. For Cursor’s MCP flaw the gap runs the other way, 8.5 from the filer and 9.8 from NVD.
A patch queue sorted on one score can put the same bug in a different week.

What Are the Top Security Risks for LLM Applications?
Prompt injection is still first on the OWASP Top 10 for LLM Applications 2026, published in August. Ranked by public incident data alone, it would drop out of the top 10.
The project leads keep it first because teams fight injection hard, so fewer clean exploits reach public databases.
- Carries 75% of the weight in the final ranking
- Ranks prompt injection #1 and misinformation near the bottom
- 7,714 incidents collected, 6,639 classified, 25% of the weight
- Puts prompt injection outside the top 10 and misinformation near the top
The order moved more than in past editions. Excessive Agency, where a model has more tools or permissions than its task needs, climbed three places because agent deployments are where damage is landing.
Unbounded Consumption rose four places. Improper Output Handling fell the furthest, from fifth to tenth.

System Prompt Leakage is gone as a name. Its replacement, Hidden Context Exposure, also covers tool schemas, retrieved policy text and developer instructions.
Several entries absorbed newer attacks: prompt injection now includes instructions hidden in images or audio, and output handling now covers the insecure code assistants generate at scale.
“Stop trying to build a model that cannot be fooled. Build the system around it, so that when the model is fooled, and it will be, nothing important breaks.”
Steve Wilson and Rock Lambros, OWASP Top 10 for LLM Applications 2026
Once a model can call tools and act on its own, OWASP points readers to a separate list. The Top 10 for Agentic Applications came out on December 9, 2025, built with more than 100 contributors.
Our AI agents statistics track how fast those deployments are spreading.
How Much Do AI-Related Breaches Cost?
A breach involving model inversion, where an attacker pulls sensitive data back out of a model, cost $6.07 million on average in the IBM Cost of a Data Breach Report 2026. That is 18% above the global average.
Ponemon Institute studied 602 breached organizations for the report, covering incidents from March 2025 to February 2026.

The two most common entry points sat around the model, not in it: compromised APIs, apps or plug-ins (27%) and cloud misconfigurations affecting AI workloads (27%).
Breaches of open-source models averaged $5.63 million, against $4.98 million for models trained in-house.
The share of organizations with an AI-related breach rose from 13% to 21%, a 61% increase by IBM’s own calculation. A year earlier, IBM’s 2025 report put that figure at 97% of organizations with an AI breach.
About one in five shadow AI incidents in 2026 ended in a regulatory fine.
Defenders use AI too. IBM found that organizations using AI and automation in security operations cut breach costs by almost $2 million on average.
SentinelOne’s guide to what artificial intelligence in cybersecurity can and cannot do weighs those gains against the risks.
Is AI-Generated Code Secure?
Across more than 100 models tested over four years, Veracode’s security pass rate has not moved: 55% in its first report and 56% in the 2026 GenAI Code Security Report. The 2026 round added 11 new models across 80 tasks.
Models now write code with correct syntax close to 100% of the time. Even so, about 44% of the coding tasks in Veracode’s tests produced a known vulnerability.

Models handle SQL injection and cryptography well and fail most cross-site scripting and log injection tasks. Java is the weakest language, with a 30% pass rate.
GPT-5.5 led the summer 2026 group at 68%, while six of the 11 new models scored between 50% and 53%.
- Coding models: models built for code averaged 51%, general-purpose models 52%.
- Reasoning: reasoning models averaged 56%, non-reasoning models 51%.
- Size: large models averaged 53%, medium and small models 51% each.
Secrets are the other leak. GitGuardian counted 28.65 million new hardcoded secrets in public GitHub commits in 2025. Keys for AI services rose 81% to 1,275,105.
Commits made with Claude Code leaked secrets at a 3.2% rate, against 1.5% across all public commits.
MCP configuration files held 24,008 unique secrets on public GitHub, 2,117 of them valid credentials. For defect rates beyond security, see our AI-generated code quality statistics and the wider AI in software development statistics.
How Do Employees Leak Data Into AI Tools?
Source code is the data employees most often paste into outside AI models, by a large margin.
Verizon’s 2026 DBIR found shadow AI became the third most common non-malicious insider action in its data loss prevention dataset in 2025, four times its share a year earlier.
Verizon counts a regular user as someone who opens an AI platform at least once every 15 days. The share signing in with personal accounts fell to 67% from 72%.
Browser extensions are a newer route out: one main job of these AI plug-ins is collecting what the user browses, internal sites included.
The best-known leak predates all of this. In April 2023 Samsung staff leaked internal data into ChatGPT, and Samsung banned generative AI tools on company devices from May 1.
In its own staff survey, about 65% said the tools carry a security risk. Our shadow AI statistics cover unapproved use in more detail.
How Are Attackers Using AI Tools?
In September 2025 a Chinese state-sponsored group used Claude Code to run an espionage campaign against about 30 targets.
Anthropic reported that AI performed 80% to 90% of the work, with humans stepping in at 4 to 6 decision points per campaign. A small number of intrusions succeeded.
The Nx attack shows how an assistant becomes the payload. While scanning for secrets, the malware tried to use installed AI tools such as Claude and Gemini.
Wiz counted over a thousand valid GitHub tokens leaked and over 5,500 private repositories made public. The Amazon Q case was luckier: no customer environments changed.
AI-assisted malware leans on well-known techniques. The median threat actor in Verizon’s data used AI help for 15 documented techniques, and under 2.5% of AI-assisted malware involved rare methods.
Phishing made up 44% of AI-assisted initial access vectors.
In the phishing simulations Verizon analyzed, the median click rate on voice and text message lures was 40% higher than on email.
IBM puts the cost side at one in four malicious breaches being AI-enabled, a 56% rise, averaging $6 million per breach. Deepfake impersonation drove 45% of those attacks.
What Do Regulators and Analysts Expect Next?
The EU pushed back its cybersecurity rules for high-risk AI. Regulation (EU) 2026/1744, published on July 24, 2026, moved obligations for Annex III systems to December 2, 2027 and for AI in products covered by Annex I to August 2, 2028.
They had been due on August 2, 2026 and August 2, 2027.
Those obligations include Article 15 of the AI Act, which lists data poisoning, model poisoning, adversarial examples and confidentiality attacks among the attacks security measures should address.
Providers of general-purpose models with systemic risk have owed “an adequate level of cybersecurity protection” since August 2, 2025, and the omnibus left that date unchanged.
Gartner’s forecasts point to more incidents. By 2028, it expects 25% of enterprise generative AI applications to have at least five minor security incidents a year, up from 9% in 2025.
It points to MCP’s design, which favors interoperability and developer speed over security enforcement by default.
“MCP was built for interoperability, ease of use and flexibility first, so security mistakes can manifest without continuous oversight for agentic AI.”
Aaron Lord, Sr. Director Analyst, Gartner, April 9, 2026
For major incidents, Gartner forecasts 15% of those applications a year by 2029, up from 3%. It also predicts that by 2029 over half of successful attacks on AI agents will exploit access control weaknesses and prompt injection.
Its survey of 316 executives ranked AI-driven discovery of vulnerabilities as the top emerging risk in the second quarter of 2026.
Budgets follow the risk. Gartner puts spending on securing AI at $2.8 billion in 2026 and nearly $7.7 billion by 2028. In its 2027 forecast, AI usage control grows fastest, at 73%.

Application security stays the largest named segment in 2027 at $851 million. Gartner expects big security vendors to buy the startups crowding into application security and usage control.
Hiring Engineers Who Can Secure AI Systems
Most of the flaws above sit in the code and infrastructure around a model: gateways, agent builders, MCP servers and CI secrets.
Second Talent matches companies with pre-vetted AI developers, AI agent developers and DevOps engineers who build and lock down these systems.
Our cybersecurity engineer rate card for Vietnam shows what the role costs. Tell us what you are building and we will send matching profiles.
Frequently Asked Questions
What is prompt injection?
It is an attack that slips instructions into text, images or files a model reads, so the model follows the attacker instead of the user. It can arrive in the chat itself or hidden in an email, web page or document the assistant processes.
What is the CISA KEV catalog?
It is the US Cybersecurity and Infrastructure Security Agency’s list of vulnerabilities with evidence of active exploitation.
US federal civilian agencies must patch listed flaws by set deadlines, and many companies use it to decide what to fix first.
Why can one CVE have two different CVSS scores?
The organization that files a CVE, often the vendor, scores it first. NVD analysts may add their own score from the same description, and they can judge attack complexity or impact differently. Both scores stay on the record.
Are AI coding assistants safe to use at work?
They have shipped serious flaws, and fixes are out for the ones listed here. The risk depends on the version you run, which MCP servers you connect and whether the agent can run commands without approval.
Treat their code like any unreviewed pull request.





