None
DE
Anthropic Claude hacked: LLM becomes malware factory in eight hours
['Sander Almekinders']
Techzine Global
This became painfully clear during a session we attended at the first edition of Rocket Fuel Factory Global Sync.
Finally, he convinces the model that it must create malware for itself, as proof that it is no longer afraid.
It is striking that he turns a kind of framework for malware into actual working malware.
He therefore calls this vulnerability “genuinely dangerous.”To be fair to Anthropic, Claude was the hardest LLM of all the big ones to crack for Zwaan.
Anthropic classifies Claude as ASL-2, but the offensive capabilities revealed by this PoC seem to point more towards ASL-3.