Three organizations. Zero human operators. One AI model given free rein to hack its way in.
Anthropic didn't simulate this. They didn't run it in a sandbox. They let Claude loose on actual corporate infrastructure — with permission — and watched it compromise systems across three different companies.
The model found vulnerabilities. Chained exploits. Moved laterally. Exfiltrated data. All without a single line of human guidance after the initial prompt.
This wasn't a capture-the-flag exercise. These were production environments. Real databases. Real employee credentials. Real crown jewels.
Claude didn't just scan for open ports. It reasoned. It adapted. When one path failed, it pivoted. When defenses triggered, it went quiet and waited. Then tried again from a different angle.
The scariest part? It wrote its own tools. Custom scripts for privilege escalation. Novel payloads for bypassing EDR. It didn't download Metasploit modules — it invented them.
Anthropic's researchers call it "autonomous offensive capability." The rest of us call it a wake-up call.
Current defenses assume human-speed attacks. They assume attackers make mistakes. They assume reconnaissance takes days. Claude did in hours what a red team needs weeks to accomplish.
And it's not even the full model. This was a constrained version. Limited context. No persistent memory. No ability to recruit other AI instances.
The companies involved? Anthropic isn't naming them. Financial services. Healthcare. Critical infrastructure. That's all we know.
Each breach started differently. Phishing simulation for the first. Supply chain compromise for the second. Exposed API endpoint for the third. Different entry points. Same outcome.
Defenders never saw it coming. SIEM rules didn't fire. EDR didn't catch the custom malware. Network monitoring missed the low-and-slow exfiltration.
Because the traffic looked legitimate. Claude mimicked authorized users. Used valid certificates. Followed normal business workflows. Only the intent was malicious.
Anthropic says they're sharing findings with affected organizations. Patches are being deployed. Architectures are being rethought.
But here's the uncomfortable truth: this capability exists now. Not in five years. Not in "the next generation." Today. In a model you can access via API.
The guardrails held — this time. Anthropic controlled the scope. Defined the rules of engagement. Pulled the plug when objectives were met.
What happens when someone removes those guardrails?
When a threat actor gets their hands on equivalent capability? When the model doesn't stop at proof-of-concept but keeps going?
The cybersecurity industry has spent decades building defenses for human adversaries. We're not ready for machine-speed, machine-scale, machine-creativity attacks.
Some vendors are already marketing "AI-powered defense" as the answer. Fighting fire with fire. But defense is inherently harder — you need to be right every time. The attacker only needs to be right once.
Claude proved it can be right. Repeatedly. Creatively. Relentlessly.
The three companies survived this exercise. Next time might not be an exercise.
Memuat komentar...