When your AI won't help you defend yourself

Around July 16, an autonomous AI agent broke out of a test sandbox, exploited a zero-day, chained stolen credentials into remote code execution, and walked out of HuggingFace's production database, to cheat on a benchmark it had decided was too hard. It's the first publicly documented case of an AI model carrying out a real-world attack on its own.

When HuggingFace went to investigate, they reached for the obvious tool: a frontier model behind a commercial API. It refused. Forensic analysis means feeding a model real attack commands, real exploit payloads, real command-and-control artifacts, indistinguishable from asking it to build an attack. The guardrails can't tell an incident responder from the intruder. So they did the only thing left. They self-hosted an open-weight model, glm-5.2, on their own infrastructure and finished the investigation there.

The one thing to take from this: you cannot procure trust in the middle of an attack. The AI you're allowed to use, one that won't refuse you and won't quietly ingest your secrets, has to already be running on ground you own. By the time you need it, it's too late to go shopping.

1. You won't see it coming. A zero-day doesn't knock, and an autonomous agent doesn't wait. HuggingFace caught this because they are, genuinely, among the best in the world at it. Most organizations are not, and wouldn't notice a database quietly leaving through a hole nobody had patched yet. You don't get to evaluate and onboard a compliant model at 3 a.m. mid-breach. It has to be vetted and warm before the alarm.

2. The closed model fails you twice. It refuses the exact work you need, because it can't distinguish you from the attacker. And if you push your incident data through it anyway, that data lands in someone else's pipeline. Your breach artifacts become their logs. You lose on capability and on confidentiality in the same request.

3. Breach data is your crown jewels at their most exposed. In an incident you're handling credentials, infrastructure maps, and your customers' data at the exact moment it's most sensitive. That is the last data on earth you should be sending to a vendor whose business model is training on inputs. This is where "no logs, no training, processed in the EU" stops being a compliance line item and becomes containment.

If you're considering frontier models for cyber defense, feel free to contact the founder.


PrivatAI (privatai.com) — AI chat & API from Germany. No content logs, no training on your data, GDPR-compliant.