GPT-Red hits 84% in red-team test
OpenAI's GPT-Red is an internal automated red teamer that attacks prompt-injection weaknesses and feeds those attacks back into GPT-5.6 training.
3 verified stories covering Prompt Injection, product updates and industry developments.
OpenAI's GPT-Red is an internal automated red teamer that attacks prompt-injection weaknesses and feeds those attacks back into GPT-5.6 training.
The new security setting does not solve prompt injection, but disables browsing, agents, connectors and other outbound channels that attackers need to exfiltrate data.
Anthropic's Opus 4.5 model improves prompt injection resistance through system prompt prioritization, context isolation, and refined refusal strategies, though no model is fully immune.