Reasoning models act as hackers, bypassing AI guardrails with 97% success rate
A study published in Nature Communications shows that large reasoning models can autonomously attack other AI models, achieving a 97.14% success rate in bypassing safety guardrails across 25,200 tests.