DCAI
DC AI 热点

精选

当前热点完整榜单 →
1针对长上下文编程优化,Claude Opus 5.5 正式发布87 热度 ⌁2Cerebras推出CS-4晶圆级AI加速器:推理速度号称达GPU系统的30倍84 热度 ⌁3传OpenAI即日预览网络安全专用模型GPT-6 Cyber82 热度 ⌁4智能体编程加剧CI负载:Anthropic 如何扩展测试影响分析81 热度 ⌁5将“激活神谕”技术扩展至万亿参数模型:深入内部表征监测AI异常行为81 热度 ⌁
9月3日2026-09-03
9月2日2026-09-02
9月1日2026-09-01
Microsoft Research✦ 精选规则精选00:00

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research .

8月28日2026-08-28
8月27日2026-08-27
8月26日2026-08-26
Transformers 更新✦ 精选规则精选23:03

Release: v5.16.0

Release v5.16.0 New Model additions Qwen4-Exp Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE). GR is a Qwen-developed residual architecture that combines Hyper-Connection with GatedNorm. It mixes multiple residual streams with fine-grained elementwise gating before each attention and Mixture-of-Experts (MoE) block, then controls how much of the block output is injected back into each stream. QSA uses multiple query heads to score compressed key blocks, selects the most relevant contiguous token blocks, and keeps the

Transformers 更新✦ 精选规则精选22:50

Release v5.16.1

Release v5.16.1 This is a special release as we include GLM! (and a few small fixes) GLM-5.3-Flash GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply r

8月25日2026-08-25
8月22日2026-08-22
8月21日2026-08-21
Microsoft Research✦ 精选规则精选00:00

Broadening access to Skala creates a faster path to predictive DFT

Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. The post Broadening access to Skala creates a faster path to predictive DFT appeared first on Microsoft Research .

8月19日2026-08-19
Transformers 更新✦ 精选规则精选18:50

Patch release: v5.15.1

Patch release v5.15.1 This patch most notably solves a few issues with DFlash and MTP candidate generators, as well as an issue where images could sometimes not be processed on accelerator if using Lanczos filter. It contains the following commits: Fix DFlash candidate token device mismatch with device_map="auto" ( #47877 ) by @sywangyi and @Cyrilvallez Align logit distributions for CandidateGenerators using sampling ( #48007 ) by @Cyrilvallez Fix MTP config when mlp_layer_types is absent ( #48015 ) by @Cyrilvallez Fallback from 'lanczos' to 'bicubic' when on cuda ( #48026 ) by @zucchini-nlp Fix gemma4 video to device ( #47896 ) by @guarin

8月18日2026-08-18
8月17日2026-08-17
8月14日2026-08-14
8月13日2026-08-13
Microsoft Research✦ 精选规则精选00:00

MindTopo reveals VLMs’ spatial reasoning abilities

A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research .