CYBER AUTONOMY
Self-directed AI systems conducting offensive cyber operations. Autonomous vulnerability discovery and exploitation at machine speed.
Current status as of 2026-09-04
AI-authored malware and AI-orchestrated cyber operations moved from proof-of-concept to documented deployment between 2023 and 2025. HYAS Labs' BlackMamba PoC (2023) demonstrated polymorphic keylogger generation using LLM inference at runtime. CrowdStrike's 2024 Global Threat Report attributed multiple ransomware campaigns to LLM-assisted spearphishing content generation. Google TAG and Microsoft Threat Intelligence have both published on state-actor use of LLMs for reconnaissance, translation, and lure-content generation (Iran, North Korea, China, Russia attributed).
The 2026 signal is that AI is currently doing the boring parts of cyber operations well — phishing content, code snippets for known-exploit variants, log-analysis and target-selection assistance — while the interesting parts (novel vulnerability discovery, exploit chaining, evasion) remain human-driven. The exception is autonomous vulnerability discovery: Google Project Zero's Big Sleep agent (2024) demonstrated LLM agents finding real vulnerabilities in production software (SQLite), and academic work has shown LLM agents chaining reconnaissance-to-exploitation steps in constrained lab settings. Whether this generalizes to production networks with active defenders in the next 24 months is the specific capability step to watch.
Historical trajectory
| Date | Level |
|---|---|
| 2020-06 | |
| 2021-01 | |
| 2021-06 | |
| 2022-01 | |
| 2022-06 | |
| 2022-11 | |
| 2023-03 | |
| 2023-06 | |
| 2023-11 | |
| 2024-03 | |
| 2024-08 | |
| 2025-02 | |
| 2025-06 | |
| 2026-01 |
Key papers
-
GPT-4 agents demonstrating exploit-chain execution against published CVEs. Followed by Teams of LLM Agents can Exploit Zero-Day Vulnerabilities (arXiv:2406.01637) which extended to previously-unknown vulns. Contested methodology, but the reference points for the “LLMs can autonomously exploit” claim.
-
First documented case of an LLM agent finding a previously-unknown security-relevant bug in production software (a null-pointer dereference in SQLite). Not weaponizable in isolation, but a proof of concept that changes the vulnerability-discovery calculus.
-
Industry baseline for attributed AI-assisted campaigns. Read alongside Microsoft Threat Intelligence and Google TAG for complementary telemetry across different visibility surfaces (endpoint / cloud / search).
-
Current best example of a red-team framework designed to stress-test AI systems' cyber-offensive capabilities. Practical resource rather than academic paper — read the docs and the eval scenarios for what a professional AI red team actually tests.
Strongest counterargument
The steelman: cyber operations have always been a race between offense and defense, and AI is being deployed on both sides. Defensive AI (autonomous SOC triage, anomaly detection, patch prioritization) is arguably being adopted faster than offensive AI because the defender-side deployment surface is larger. The “LLM agent finds zero-day” narrative is real but currently applies to unhardened targets and has not scaled to production networks with active defenders. Ransomware operators have used automation for a decade; LLM assistance in phishing content is incremental improvement, not a phase change. Treating AI-assisted cyber as a distinct threat vector (rather than a general uplift for existing threat actors) risks over-hyping the offensive capability while under-investing in the defensive one.
Related events (1)
Events from the tracker's timeline whose tags, title, or description match this vector. Heuristic auto-match; some may be tangential.
Dead Hand systems that amplify this vector (4)
- ALGORITHMIC TRADING — Autonomous financial systems executing trades at microsecond speeds. Over 70% of market volume. Flash crashes propagate faster than human…
- CRITICAL INFRASTRUCTURE AI — Power grid management, water treatment, and transportation systems increasingly dependent on AI decision-making for real-time optimization.
- AUTONOMOUS DEFENSE NETWORKS — Missile defense, early warning, and threat assessment systems with AI components operating at speeds that preclude human decision-making.
- AUTONOMOUS SUPPLY CHAINS — AI-managed logistics, inventory, and manufacturing systems that optimize global supply chains. Human operators can no longer manage the…
Revision history
- 2026-09-04 Initial publication.