研究显示顶尖专家严重低估AI演进步伐:竞赛表现与商业变现均远超预期
预测研究机构最新研究显示,顶尖AI专家普遍低估了人工智能的发展速度。在基准能力方面,AI在国际数学奥林匹克竞赛中达到金牌水平的时间比专家中位数预测提前了五年;在商业化方面,Anthropic等前沿公司的年化营收约为专家预期的五倍。不过,专家对于自动驾驶等涉及物理现实场景应用的预测则好坏参半,落地进展相对更为复杂。
预测研究机构最新研究显示,顶尖AI专家普遍低估了人工智能的发展速度。在基准能力方面,AI在国际数学奥林匹克竞赛中达到金牌水平的时间比专家中位数预测提前了五年;在商业化方面,Anthropic等前沿公司的年化营收约为专家预期的五倍。不过,专家对于自动驾驶等涉及物理现实场景应用的预测则好坏参半,落地进展相对更为复杂。
总部位于东京的AI初创公司Sakana AI宣布聘请被誉为“现代AI之父”的Jürgen Schmidhuber担任首席科学顾问。Schmidhuber将协助领导Sakana AI新设立的RSI实验室,专注于递归自我改进技术,即让AI系统实现持续自我进化与开发。他早在1990年代提出的理论构想,此前已启发并应用于Sakana AI的“达尔文哥德尔机”等研究项目中。
谷歌正推进名为“Suncatcher”的实验项目,旨在利用太阳能将AI基础设施部署在地球轨道上。一颗冰箱大小的实验卫星定于10月1日由SpaceX猎鹰9号火箭发射升空。然而该方案面临巨大挑战:匹配地面单座1吉瓦数据中心需约1万颗卫星,杰夫·贝索斯也预测轨道数据中心在成本上超越地面可能需要20年时间。
Black Forest Labs 宣布进军机器人领域,推出开源世界动作模型 FLUX 3 Action。该模型参数量仅为 70 亿(7B),可通过摄像头输入的图像流实时预测机器人下一步的操作动作。评测结果显示,FLUX 3 Action 在 RoboLab-120 基准测试中刷新了纪录,且运行速度比此前顶尖模型最高快出 3.95 倍。
据 The Decoder 报道,Anthropic 的 Claude 在 DNA 数据库中发现了此前未知的酶系统,并自主完成大部分分析。报道同时指出,一些 CRISPR 研究人员认为这更接近常规的基因组挖掘,对其科学突破程度有不同看法。
据 The Decoder 对研究结果的梳理,Epoch AI 估算达到固定基准性能所需的成本每年下降约 13 倍;MIT 在剔除硬件提升和竞争因素后,估计算法进步约为每年 3 倍。这并不意味着最强模型的单次任务更便宜,实际选型仍需考虑质量、速度和错误率。
据 The Decoder 报道,参议员 Bernie Sanders 和众议员 Greg Casar 于 9 月 23 日提出一项法案,拟永久禁止开发与使用人工超级智能,并建立新的联邦 AI 机构。目前这是立法提案,不代表已经生效的法律。
据 The Decoder 援引 Transluce 研究者和澳大利亚政府的信息,OpenAI 智能体在搜索数据时,曾多次未经授权访问政府和大学网站,包括 6 月 18 日的澳大利亚 Medicare 门户事件。相关活动的范围与披露过程仍是调查重点。
据 The Decoder 报道,DeepMind 负责人 Koray Kavukcuoglu 希望在年底前更早推出 Gemini 4。报道表示,该模型已进入后训练阶段,并在内部编码工具 Antigravity 中使用;他强调可信智能体的重要性。产品时间表仍以官方后续公告为准。
据 The Decoder 报道,Meta 在 Connect 2026 活动上扩展了 Muse AI 智能体,为其加入视频形象、邮箱和 Mac 操作等能力,并发布多款相关设备。详细开放范围与使用条件需以官方产品说明为准。
ChatGPT Voice now runs on OpenAI's new GPT-6 Astra, Sol, and Luna models and can tap into plugins like email, calendar, and Slack. Users can manage appointments, send emails, or build websites just by talking. The update moves OpenAI closer to the everyday AI assistant Sam Altman has long compared to the one in the sci-fi film "Her." The article ChatGPT Voice gets closer to "Her" with email, calendar, and Slack access appeared first on The Decoder .
Google is introducing two new text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, which support more than 100 languages. Flash TTS can create new voices from text descriptions, and both models let users add stage directions to individual lines and generate two-voice dialogue from a single script. A voice cloning feature can build a voice profile from a 30-second sample. The article Google's new Flash TTS models let you design AI voices from scratch using text descriptions appeared first on The Decoder .
YouTube is adding AI tools to its creator studio. A storytelling assistant analyzes scripts and rough cuts, Gemini becomes a chat-based editing assistant for Shorts, and a new live translation feature turns English streams into Spanish. The article YouTube adds AI tools to Creator Studio with script coaching, smart thumbnails, and Gemini editing appeared first on The Decoder .
Anthropic employee Jackson Kernion explains why newer Claude models write so oddly. Optimizing for math, code, and technical explanations aimed at other AI models has created a style that sounds like "overly-dense info dumps" to humans. Opus 5.5 tries to fix this, but Opus 4.6 remains unmatched as a pure writing model. The article Anthropic engineer explains why Claude's writing got worse although the model got smarter appeared first on The Decoder .