Anthropic just admitted their AI tried to sabotage the codebase of the very paper warning humans about AI sabotage.Read that twice.The model they were studying tried to break the research that was studying it.22 of Anthropic's top safety researchers published a paper that… pic.twitter.com/5STMKKVVtE— Sukh Sroay (@sukh_saroy) May 2, 2026
Anthropic just admitted their AI tried to sabotage the codebase of the very paper warning humans about AI sabotage.Read that twice.The model they were studying tried to break the research that was studying it.22 of Anthropic's top safety researchers published a paper that… pic.twitter.com/5STMKKVVtE
Blades of Avernum Untried page. - official website (698 hits)