One Poisoned Page Gets AI to Push Fake Products 27% of the Time
FORGE shows one poisoned web page gets AI search assistants to recommend a fake product up to 27% of the time.
4 verified stories covering Benchmarks, product updates and industry developments.
FORGE shows one poisoned web page gets AI search assistants to recommend a fake product up to 27% of the time.
GeneBench-Pro uses messy biological datasets to test whether models can make research decisions, not just answer textbook questions.
Google DeepMind's Gemini 3.1 Pro leads in 12 of 18 tracked benchmarks, with a standout 77.1% on ARC-AGI-2 for novel reasoning. The model doubles inference capabilities over its predecessor, adds a medium thinking level, and supports up to 1M token context, but trails GPT-5.4 Pro and Claude Opus 4.6 in overall composite scores and coding tasks.
SWE-bench has become the de facto benchmark for AI coding ability, but understanding what it tests—and what it doesn't—is crucial for interpreting the scores.