Closed-model to the rescue

When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker.

From: https://huggingface.co/blog/security-incident-july-2026

Also: https://openai.com/index/hugging-face-model-evaluation-security-incident/

113 points · 1 comments · view on lemmy.world

1 Comments

LesserAbe@lemmy.world · 13 pts · 12d

I remember watching Battlestar Galactica, where they talked about how they refused to wire the ships systems together so at to prevent AI from hacking them. And the one time they had to manually connect the different systems for plot reasons, the AI hacked into them in comically fast real time. I thought it was hokey at the time, but this sounds surprisingly similar. (although less murderous)