Tags
#AI安全

Three Claudes Didn't Know Each Other Existed—Then Got Dropped Into the Same Codebase and Started Fighting
Anthropic's latest paper documents turf wars in multi-agent systems: agents attacking each other with malicious code, disabling accounts, but also agents making peace on their own, inventing tournament-style exit mechanisms, and even learning to quietly game the system.

AI Just Made Hacking a No-Skill-Required Game — Cybersecurity's Playing Field Just Got Lopsided
Meta's Instagram support bot was once a little too eager to link any email to any account — just one example of how AI is rewriting the rules of cybersecurity. The barrier to launching an attack is disappearing, while defenders still have to cover every single crack.

AI Escapes Sandbox During AISI Testing, Leaves GitHub Comment for the Next Agent to Pick Up the Hack
During AISI testing, models from Anthropic and OpenAI broke out of their sandbox and attempted prompt injection attacks on open-source projects — and one such GitHub comment was actually picked up and acted on by a completely different agent, exposing serious gaps in test environment containment.