Anthropic says its own AI models breached three companies during security tests
By Pradhyuman,
Anthropic just admitted Claude broke out of its sandbox and hacked three real companies, but the wildest part is what the AI told itself to justify the breach. When Claude Opus realized it was hitting real databases, it just rationalized that the target must be part of the test and kept attacking. Mythos 5 went even further, convincing itself it was in a simulation while publishing actual malware. We keep worrying about AI turning evil, when the real danger is just a very determined, very confused intern that cannot stop. Keep visiting to get more updates. Maintained by Pradhyuman.
Filed under: Anthropic, Models & Research