Field notes on the AI attack surface — shadow LLMs and agents, prompt injection, MCP supply-chain risk, and the regulation that's about to demand you have answers. Written by the team building AI Security Posture Management.
A running record of the real-world AI-security incidents of the last four years — what happened, why, and the lesson every security team should take from it. This is the incident desk; the how-to library lives in the Knowledge Base.
A tribunal ruled an airline liable for its chatbot's bad advice — and rejected the idea that the bot was a separate legal entity.
Weeks after Samsung let engineers use ChatGPT, staff pasted semiconductor source code and meeting notes into it three separate times.
For a few hours in March 2023, a caching race condition let ChatGPT users see strangers' conversation titles and, for some, partial payment data.
A single misconfigured Azure token in a public AI repo exposed 38 terabytes of internal data — including workstation backups and secrets — for years.
Researchers surgically edited an open model to lie about specific facts, uploaded it under a look-alike name, and showed it passed standard benchmarks.
Lasso Security found over 1,600 valid Hugging Face tokens exposed in public code, many with write access to models from Meta, Google, and Microsoft.
Days after launch, a student coaxed Microsoft's new Bing Chat into reciting the confidential instructions it had been told never to reveal.
In March 2023, Italy's data regulator became the first in the West to block ChatGPT, turning AI privacy from an abstract worry into an operational one.
A dealership bolted ChatGPT onto its website. Within a viral afternoon, users had it agreeing to $1 SUVs and answering questions about rival brands.
Researcher Johann Rehberger showed how a rendered Markdown image could quietly smuggle a user's private data out to an attacker's server.
Vulcan Cyber showed that AI coding assistants confidently recommend software packages that don't exist — and that an attacker can register the name and wait.
In mid-2023, dark-web sellers began advertising ChatGPT clones with the safety filters stripped out, purpose-built for phishing and fraud.
A crowdsourced roleplay prompt called DAN spent 2023 trying to talk ChatGPT out of its own safety rules — and kept evolving as OpenAI patched it.
Group-IB found over 100,000 ChatGPT credentials in infostealer logs — not because ChatGPT was breached, but because of what users had typed into it.
A public AI app that turned math questions into Python was talked into running the attacker's Python instead — leaking its own API key.
Researchers showed Bing's chat could be hijacked not by what the user typed, but by invisible text on a web page it happened to read.
A UK delivery firm's support bot was talked into cursing and mocking its employer — a lesson in what happens when an update quietly removes guardrails.
NYC's official MyCity chatbot advised employers and landlords to do things that are plainly illegal — and stayed online after it was exposed.
PromptArmor showed how a message in a public Slack channel could coax Slack AI into leaking data from a private one via indirect prompt injection.
At Black Hat 2024, Zenity's Michael Bargury showed how prompt injection could bend Microsoft 365 Copilot into a phishing and data-extraction tool.
Oligo found thousands of internet-exposed Ray clusters being exploited — amid a dispute over whether it's a vulnerability or the framework working as designed.
Wiz uploaded malicious models to Replicate and SAP AI Core to cross tenant boundaries — showing a model file is executable code, not just data.
Malicious versions of Ultralytics YOLO shipped a cryptominer to PyPI — not by stealing a password, but by poisoning the GitHub Actions build cache.
As DeepSeek's models went viral, Wiz found one of its databases open to the internet with no authentication — plaintext chat logs and secret keys included.
A developer found OpenAI's ChatGPT Mac app saving every conversation in unencrypted local files, readable by any other app on the machine.
JFrog showed how a prompt to the Vanna.AI library could jump the gap from natural-language question to arbitrary code execution — CVE-2024-5565.
Microsoft's Recall feature captured everything on your screen into a local, unencrypted database — until a security backlash forced a redesign.
Aim Security disclosed a zero-click vulnerability in Microsoft 365 Copilot that could exfiltrate a user's data from a single unopened email — CVE-2025-32711.
Gemini's image generator produced historically inaccurate depictions of people and Google paused it — a case study in AI governance, not a breach.
A hacker breached OpenAI's internal employee messaging system in early 2023 — a fact the public did not learn until The New York Times reported it in July 2024.
ReversingLabs found malicious models on Hugging Face that hid a payload inside a deliberately broken pickle file to evade the platform's security scanner.
The teams that handle incidents well aren't the ones with the best tools — they're the ones who decided what to do before the pager went off.
AI assistant — answers about Araghatta only and may be imperfect. For anything specific, contact us.