DCAI
DC AI 热点

全部 AI 动态

9月25日2026-09-25
Google ResearchAI 评分 45/10003:40

自动化连贯长视频生成技术研究

该内容聚焦于利用生成式人工智能技术实现连贯长视频的自动化生成。长视频生成中的时序连贯性和自动化一直是视频生成领域的重要挑战。不过,由于当前提供的原始素材信息极度简略,仅提及生成式人工智能与长视频生成主题,未包含具体的模型架构、算法实现细节或实验评测数据,具体的技术突破与实现路径仍需结合后续完整资料进一步了解与评估。

阅读原文 ↗推荐理由:长视频生成的连贯性是当前多模态研究热点,但本条素材提供的信息量严重不足。# 视频生成# 长视频# 生成式AI# 时序连贯性# 多模态
Google DeepMind✦ 精选AI 评分 65/10000:20

Gemini 3.8 Live 发布,支持实时数字人化身功能

根据最新消息,Gemini 3.8 Live 正式推出,并引入了实时化身(Live Avatar)交互功能。该功能预计将为多模态实时交互带来具有视觉形象的数字人体验。由于当前输入素材仅包含标题,缺乏具体的技术规格、功能演示及发布细节等详细正文内容,具体能力表现与上线范围仍需等待官方进一步披露。

阅读原文 ↗推荐理由:涉及 Gemini 系列的实时交互与数字人新功能,但因缺乏详细正文信息,推荐度适中。# Gemini# 多模态# 实时交互# 数字人# AI助手
9月24日2026-09-24
Hugging FaceAI 评分 45/10022:08

借助 LFM2.5-VL-DSpark 加速视觉语言模型

该内容主要探讨如何通过 LFM2.5-VL-DSpark 针对视觉语言模型(VLM)进行性能加速与推理优化。由于提供的原始材料仅包含标题,缺乏详细的摘要和正文内容,目前关于其具体的加速架构、基准测试表现、计算效率提升幅度及技术实现细节均无法获知,需等待官方或作者补充更详尽的技术报告与文档说明。

阅读原文 ↗推荐理由:涉及多模态模型加速优化方向,但因缺乏详细上下文,建议等待完整技术报告披露。# 视觉语言模型# 模型加速# 多模态# 推理优化# AI模型
arXiv 自然语言处理✦ 精选AI 评分 75/10012:00

引导与动作的相互强化:多模态网页智能体离线评测基准 WebMRE

针对网页智能体在实时环境中评估存在状态漂移、难以稳定复现和开展可控训练研究的痛点,研究团队提出了离线评估基准 WebMRE。该基准基于 WebArena 的成功轨迹构建,包含 541 个任务和 5,293 个步骤,配备经过全面审计的测试标签和确定性评估协议,无需运行环境即可确保同一检查点在多次运行中输出一致评分。每个步骤将人类指导语句与具体动作配对,首次实现了对引导与动作相互强化机制的离线研究。

阅读原文 ↗推荐理由:提出了无需实时环境的确定性离线基准,有效解决了网页智能体评测难以复现和量化的核心痛点。# 网页智能体# 离线基准# WebArena# 智能体评估# 多模态
arXiv 人工智能✦ 精选AI 评分 75/10012:00

ShowTellArena:评估智能体从业务演示中理解工作流能力的基准数据集

研究人员提出了 ShowTellArena,这是一个用于评估智能体在观看带有口述解说的业务演示后其理解能力的基准测试协议与公开数据集。该基准模拟了人类通过“边演示边讲解”指导新同事的场景。v1.0 版本涵盖财务、招聘、采购、客户决策、库存和物流等领域的 50 个业务工作流任务,包含屏幕录像、截图、旁白解说及 502 道评估试题,旨在系统性测试智能体对复杂商业流程的掌握水平。

阅读原文 ↗推荐理由:针对智能体从多模态演示与讲解中学习业务工作流的能力,提出了标准化的评测协议与多行业数据集。# 智能体# 基准测试# 工作流自动化# 多模态# ShowTellArena
arXiv 自然语言处理✦ 精选AI 评分 70/10012:00

PRISM-VLM:面向紧凑型视觉语言模型的多维度判别评测基准

针对现有评测基准多沿袭前沿大模型设计、导致紧凑型视觉语言模型难以拉开区分度的问题,研究人员推出了 PRISM-VLM 评测基准。该基准围绕轻量级多模态模型的常见失效模式,从任务质量、行为鲁棒性和能力瓶颈等七个关键维度对每个评测样本进行综合打分,并将其融合为单一的 PScore 指标,从而更精确、精细地评估紧凑型多模态模型的真实表现与局限性。

阅读原文 ↗推荐理由:针对轻量级多模态模型评测区分度不足的痛点,提出了涵盖多失效模式的细粒度评测体系。# 视觉语言模型# 基准评测# 多模态# 紧凑模型# 模型评估
9月3日2026-09-03
8月26日2026-08-26
Transformers 更新✦ 精选规则精选23:03

Release: v5.16.0

Release v5.16.0 New Model additions Qwen4-Exp Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE). GR is a Qwen-developed residual architecture that combines Hyper-Connection with GatedNorm. It mixes multiple residual streams with fine-grained elementwise gating before each attention and Mixture-of-Experts (MoE) block, then controls how much of the block output is injected back into each stream. QSA uses multiple query heads to score compressed key blocks, selects the most relevant contiguous token blocks, and keeps the

Transformers 更新✦ 精选规则精选22:50

Release v5.16.1

Release v5.16.1 This is a special release as we include GLM! (and a few small fixes) GLM-5.3-Flash GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply r

8月10日2026-08-10
Transformers 更新✦ 精选规则精选18:28

Release: v5.15.0

Release v5.15.0 New Model additions Meta Muse Glimmer Muse Glimmer, released today, is Meta’s new multimodal model, especially designed for agentic use cases. Distilled from Muse to 30B parameters, and released under the Apache 2.0 license, it can be deployed to local setups for privacy-aware applications such as coding, document analysis, personal assistants, Claw- or Hermes-like setups. Muse Glimmer is a dense 30B parameter model consisting of: 2B ViT-style encoder for vision (Perception Encoder) 28B parameter text decoder We're covering it in the following blogpost: http://hf.co/blog/muse-glimmer GraniteMoeSWA & GraniteSWA Links: Documenta

7月16日2026-07-16
Transformers 更新✦ 精选规则精选03:02

Release v5.14.0

Release v5.14.0 New Model additions Inkling (fresh from Thinking Machines): 975B total, 41B active Add Inkling model #47347 by @molbap @Cyrilvallez @eustlb and @zucchini-nlp Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI- powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-foll

7月4日2026-07-04
Transformers 更新✦ 精选规则精选00:06

Release v5.13.0

Release v5.13.0 New Model additions KimiK 2.5, 2.6, and 2.7 This release includes the architecture for Kimi 2.5 which is used by 2.5-2.7: Kimi K2.5 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration. The model was proposed in Kimi K2.5: Visual Agentic Intelligence and further improved in [Kimi K2.6: Advancing Open-Source Coding](Kimi K2.5: Visual Agentic Intelligence). Kimi K2.5 achieves significant improvements on complex, end-to-end coding tasks, generalizing robustly across programming langua

6月9日2026-06-09