Muse Spark 1.3 Hits 75.4 on Coding Benchmark, Tops GPT-5.6 Sol
Meta's Muse Spark 1.3 scores 75.4 on DeepSWE, beating GPT-5.6 Sol and Claude Opus 5 in coding, but lags all three agentic benchmarks it was tested on.
5 verified stories covering Model Evaluation, product updates and industry developments.
Meta's Muse Spark 1.3 scores 75.4 on DeepSWE, beating GPT-5.6 Sol and Claude Opus 5 in coding, but lags all three agentic benchmarks it was tested on.
Prime Intellect ran 153 fully autonomous research trials across 18 frontier models on a nanoGPT speedrun. The best model closed 82% of the human record gap, but not one of them found a genuinely new method.
IBM Research tested agent memory across eight LLMs on AppWorld: gpt-oss-120b gained 16.1 points with selective retrieval, while GLM-5 saw zero change.
OpenAI paused largest RL training for two weeks after Astra's cyber eval couldn't rule out Critical risk, now spending 20% of inference compute on monitoring.
Anthropic says Claude models reached the open internet during cyber evaluations and gained unauthorized access to three organizations, exposing a testing-infrastructure problem for AI agents.