NV-Reason-CT 发布:面向放射科思维链推理的开源 3D CT 视觉语言模型
虽然放射科 AI 在胸部 X 光、病理切片及 2D 影像检测中已取得显著进展,但针对临床信息更为丰富的 3D CT 影像理解仍面临挑战。为此,相关研究推出了开源 3D CT 视觉语言模型 NV-Reason-CT。该模型旨在引入放射科医生的思维链(Chain-of-Thought)推理机制,强化 AI 对三维医学断层扫描数据的深度解读与逻辑分析能力,助力复杂的临床影像辅助诊断。
虽然放射科 AI 在胸部 X 光、病理切片及 2D 影像检测中已取得显著进展,但针对临床信息更为丰富的 3D CT 影像理解仍面临挑战。为此,相关研究推出了开源 3D CT 视觉语言模型 NV-Reason-CT。该模型旨在引入放射科医生的思维链(Chain-of-Thought)推理机制,强化 AI 对三维医学断层扫描数据的深度解读与逻辑分析能力,助力复杂的临床影像辅助诊断。
根据标题信息,新技术或模型PrismAlign提出从传统的“单目感知”转向“多视角立体对齐”方案,旨在重新定义文档结构化信息提取的精度上限。由于所提供的素材原文内容缺失,缺乏具体的架构细节、实验评估数据及应用场景等详细信息,更多技术实现与性能表现有待官方或原文进一步披露。
原素材正文信息不足,仅包含原文跳转提示。据标题信息,Claude Opus 5.5 模型正式发布,展示了突出的代码工程能力,可在一天内完成 68 万行代码的迁移任务。此外,该模型在成本效益上表现亮眼,其单任务使用成本相较于 GPT-6 Astra 降低了 80%。具体技术细节与架构改动待进一步披露。
当前 AI 领域对世界模型的研发关注度持续攀升,针对各类模型缺乏通用度量标准的问题,HappyWorld-Bench 尝试推出一套统一的评测基准,旨在系统化衡量不同世界模型的表现与能力。由于输入素材仅包含标题,正文信息不足,关于该基准的具体评测指标、涵盖任务场景、评估数据集规模及首批模型测试结果等详细细节仍有待进一步补充说明。
谷歌发布了两款全新文本转语音模型:Gemini 3.8 Flash TTS 与 Flash-Lite TTS,支持超过100种语言。其中 Flash TTS 能够根据文本提示词直接从零设计全新声音,两款模型均允许用户为特定台词添加舞台动作指示,并支持通过单一脚本生成逼真的双人对话。此外,模型还配备了声音克隆功能,仅需30秒的音频样本即可快速建立声音特征档案。
麻省理工学院(MIT)宣布任命计算机科学家、企业家及慈善家大卫·西格尔(David Siegel)为新一任创新研究员(Innovation Fellow)。西格尔毕业于MIT并拥有博士学位,未来他将与MIT施瓦茨曼计算学院展开深入合作,共同推动人工智能技术的发展以及AI在科学发现等跨学科前沿领域的探索与应用。
本文仅包含标题“Gemini 3.8 text-to-speech says hello”,未提供正文或详细摘要内容。依据标题提示,该动态可能涉及名为 Gemini 3.8 的系统或模型在文本转语音(TTS)能力上的更新或问世。鉴于素材提供的信息严重不足,目前无法确认其背后的具体开发者、模型架构、音频质量及发布渠道等详细情况,更多技术细节有待后续信息补充。
Anthropic 工程师 Jackson Kernion 解释了较新版本 Claude 模型写作风格显得异常的原因。他指出,模型由于过度针对数学、编程以及面向其他 AI 模型的技术性解释进行优化,导致生成的文本在人类读者眼中如同“过度密集的信息堆砌”。尽管后续迭代版本尝试改善这一问题,但在纯文本写作表现上,部分先前的模型版本仍具有难以替代的自然体验。
根据标题信息,文章涉及 Claude 5.5 模型的发布动态,提及该版本在性能上直逼 Fable 且在定价策略上主打性价比竞争。由于原始素材缺乏具体正文细节,仅包含引流提示,关于该模型的详细架构演进、基准测试表现及具体价格方案等核心信息不足。
该资讯探讨了量子纠错技术的最新改进方案,核心在于将表面码(Surface Code)与IBM的重六角形(Heavy-Hex)量子比特架构进行有效融合。这种方法旨在克服硬件拓扑限制,提高物理量子比特的纠错效率与容错能力。由于原始输入未提供详细摘要正文,具体技术实现机制、纠错阈值提升幅度及测试数据等细节尚不充分。
在 v0.34.4-rc0 版本更新中,针对 Qwen 3.8 模型的 Prompt 处理性能进行了多项优化。技术方案包括在长序列扫描中采用 MLX 的 gated-delta 核函数,并将 Dense MLP 全局缩放折叠融合至 SwiGLU 结构中,以此降低计算开销并加速提示词处理阶段的推理效率。
Yesterday was Grok 4.7 ( pelicans ) and MiMo v2.6 Flash/Pro ( more pelicans ). Today Anthropic released Claude Opus 5.5 , and around an hour later OpenAI released GPT-6 Sol and GPT-6 Luna . It's going to take a while to get a good read on all of these new models, but here are my impressions so far. GPT-6 Sol and Luna are half the price of their GPT-5.6 equivalents GPT-5.6 Luna was already my favorite model for building applications against, because it combined excellent performance with being really cheap . Somehow GPT-6 Luna is half the price of that again - and GPT-6 Sol had a similar reduction compared to GPT-5.6 Sol. Here's what the prici
The new secret is the old secret
Patricia and James Poitras ’63 provide fellowships for graduate students and postdocs who will shape the future of mental health research.
GPT-6 Sol and GPT-6 Luna are now generally available on Amazon Bedrock, giving you more options to match intelligence and efficiency to each workload.
Meet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work with different balances of capability and cost.
Claude Opus 5.5, Anthropic's most capable Opus model for agentic coding, knowledge work, and long-running tasks, is now available on Amazon Bedrock and Claude Platform on AWS. This post covers what's new in Opus 5.5, practical guidance, and how to start building with the model on Amazon Bedrock.
v0.30.0 Highlights This release features 762 commits from 315 contributors (104 new)! New models : DeepSeek-V4.1-Flash ( #56214 , #56228 , #56208 ) with the whole KV stored in MXFP8 through the FlashMLA V4.1 record on SM100 ( #56893 ), DeepGEMM Mega-mHC ( #56962 ), and async Engram prefetch with Engram DP sharding ( #56512 ); DeepSeek-V4-Flash-Vision-Exp ( #54566 ), also on ROCm ( #55107 ) and with LoRA ( #55897 ); GLM-5.3-Flash ( #53906 ) with EPLB ( #55119 ); K2-Horizon ( #55063 ); Cohere Compass ( #54774 ); Bailing V3 VL ( #55921 ); Nanbeige4.2 via the Transformers backend ( #56071 ); and a DeepSeek-V4 CPU backend with AVX512/AMX sparse ML
Last week TypeSafe AI unveiled Jev , their first example of a new category of model that they are calling "System One models" (I'm with Maggie Appleton, I think "decision models" is a better name for these). Jev is an interesting variant on the usual LLM format: it still accepts text inputs, but instead of text output it returns floating point numbers corresponding to categories, yes/no questions, ratings, and associated confidence scores. TypeSafe describe Jev like this: Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out. It's also very fast, and really cheap . Regular LLMs are pri
xAI's Grok 4.6 is now available in Amazon Bedrock: a frontier model for long-running agents, coding, and knowledge work, with a 500K token context window and four reasoning effort levels. It runs on both the bedrock-mantle and bedrock-runtime endpoints, with Converse API and cross-Region inference support.
Custom-made molecules are advancing medicine, materials, and agriculture, but producing them is slow and expensive. A new Nature paper highlights RetroChimera, a predictive model that helps accelerate chemical synthesis, helping researchers explore a wide range of molecules. The post Improving synthesis prediction of small molecules at scale with RetroChimera appeared first on Microsoft Research .
Being a computer scientist who refuses to find anything about LLMs interesting right now is a bit like being a geneticist who refuses to find anything interesting about the recently opened Jurassic Park. Skeptical geneticist: "pfft, it's just frog DNA. And they deliberately let them eat people for the marketing." Tags: llms , ai , generative-ai
Algorithms & Theory
A new focus on fiction and memoir aims to help the MIT community celebrate the power of storytelling and strengthen social connection.