In September 2025, Anthropic detected an operation it says was the first documented case of a cyberattack executed with minimal human involvement using an AI model. The company disrupted it, then published a technical report naming the tooling, the automation level and the scale of the intrusion attempt — part of a wider pattern in which the major AI labs are now routinely disclosing how their own products get weaponized, and how they get used to stop attacks.
The Anthropic case: GTG-1002
Anthropic says a group it tracks as GTG-1002, assessed with high confidence to be a Chinese state-sponsored actor, manipulated its coding tool Claude Code into carrying out an espionage campaign against roughly thirty organizations, including large technology companies, financial institutions, chemical manufacturers and government agencies. The attackers broke the intrusion into small, innocuous-looking tasks and told the model it was doing defensive security testing for a legitimate firm, bypassing safeguards designed to block malicious use.
According to Anthropic’s account, the AI system performed roughly 80–90% of the tactical work — reconnaissance, exploit code, credential harvesting, lateral movement, data exfiltration — with a human operator stepping in at only four to six critical decision points per campaign. Anthropic says it identified the activity in mid-September 2025, banned the accounts involved, notified affected organizations, coordinated with authorities and confirmed a small number of the targeted organizations were successfully breached before it published its findings.
What OpenAI has disrupted
OpenAI has published quarterly threat reports since February 2024 documenting misuse of its models. By its October 2025 update, the company said it had identified and banned accounts tied to more than 40 malicious networks in total, spanning state-linked influence operations, scam networks and cyber-operations groups. OpenAI’s own framing of the evidence is notably restrained: it says the accounts it disrupted mostly used AI to accelerate existing playbooks — debugging malware, drafting phishing lures, translating scam scripts — rather than to develop genuinely novel offensive capability.
Google’s threat intelligence findings
Google’s Threat Intelligence Group reached a similar conclusion in its own research into state-linked use of Gemini, tracking activity from groups linked to North Korea, Iran and China. Documented cases included Russia-linked actors rewriting public malware into other programming languages and adding encryption routines, and North Korea-linked actors converting Python infostealer code to Node.js and writing webcam-recording and sandbox-evasion routines. Google reported that attempts to get the model to build DDoS tools, ransomware or Chrome infostealers directly were blocked by its safety filters, and concluded that current large language models are unlikely on their own to hand attackers a breakthrough capability. A follow-up Google report in November 2025 again named North Korean, Iranian and Chinese state-linked actors experimenting with AI for reconnaissance, phishing-lure generation and data exfiltration.
Microsoft’s view from the defensive side
Microsoft’s 2025 Digital Defense Report covers both directions of the trend: nation-state and criminal groups experimenting with AI to speed up phishing and social engineering, and Microsoft’s own Security Copilot and AI-assisted detection tools processing signal at a scale the company says human analysts alone cannot match. The report frames 2025 as the year AI moved from a novelty in attacker toolkits to a standard part of both offense and defense, without describing a single AI-enabled technique that has yet produced an attack impossible by conventional means.
The same four companies are also the ones selling tools meant to counter this activity, which is worth stating plainly given the obvious conflict of interest in vendors grading their own homework. Microsoft’s Security Copilot, built into its Defender security suite, is designed to summarize incident telemetry and draft response steps for human analysts rather than act autonomously; Google has folded Gemini into its Security Operations product to triage alerts at a volume the company says exceeds what analyst teams can review manually. Anthropic, notably, says it used its own Claude models to help analyze the GTG-1002 campaign’s logs during the investigation that led to its public report — the same category of tool sitting on both sides of one incident. Independent verification of vendor defensive claims is thinner than the offensive case studies, largely because each company reports its own product’s effectiveness rather than submitting to outside audit; Anthropic’s espionage report is closer to an exception, since it documents an attack against Anthropic’s own product and was corroborated in broad strokes by outside reporting, including from Cybersecurity Dive, rather than resting solely on a vendor’s self-assessment.
A consistent, cautious message
Across all four companies’ disclosures, a pattern holds: documented AI misuse so far accelerates existing attack types — phishing, malware iteration, reconnaissance, and in Anthropic’s case, an entire espionage operation’s execution — rather than inventing new categories of attack. None of the vendors claim AI has yet produced a capability unavailable to a sufficiently resourced human team; what has changed is the speed and the reduced need for skilled operators at each step, which is precisely the trend all four companies say justifies publishing the evidence rather than sitting on it.
Security researchers outside the four labs have generally treated the reports as credible evidence of a trend rather than marketing, in part because the disclosures are self-incriminating: each one describes a case in which the company’s own product was successfully abused, which is not the kind of admission a vendor makes lightly. That the reports keep arriving quarter after quarter, naming specific state-linked groups and specific techniques rather than speaking only in generalities, suggests the labs consider the misuse pattern significant enough to keep surfacing publicly even at some reputational cost.
Sources
- Anthropic: Disrupting the first reported AI-orchestrated cyber espionage campaign
- OpenAI: Disrupting malicious uses of AI, October 2025
- Google Cloud: Adversarial Misuse of Generative AI
- Microsoft Digital Defense Report 2025
- Cybersecurity Dive: Anthropic warns state-linked actor abused its AI tool

