ASR Models Recite Numbers That Were Never Spoken
A Hume AI probe of 11 leading speech recognition models finds the top scorers on public benchmarks are also the most likely to echo wrong reference transcripts instead of what they actually hear.
6 verified stories covering AI benchmarks, product updates and industry developments.
A Hume AI probe of 11 leading speech recognition models finds the top scorers on public benchmarks are also the most likely to echo wrong reference transcripts instead of what they actually hear.
OpenAI says GPT-5.6 Sol’s ARC-AGI-3 public-set score rose from 13.3% to 38.3% after retaining reasoning and enabling compaction.
RoboDojo shows real robot success is still low is reframed for global technology readers, with the key numbers, product claims and open questions kept intact.
Alibaba tests agents with HS code classification is reframed for global technology readers, with the key numbers, product claims and open questions kept intact.
The 750-task benchmark shows leading models can help with life science, but still fall far short of independent research work.
A Shanghai Jiao Tong University-led benchmark shows agents often find the right files but identify the exact bug-relevant lines only about 14% to 19% of the time.