DCAI
DC AI 热点

全部 AI 动态

9月1日2026-09-01
Microsoft Research✦ 精选规则精选00:00

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research .

8月28日2026-08-28
MIT News · AI规则精选03:20

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

8月27日2026-08-27
8月26日2026-08-26
Transformers 更新✦ 精选规则精选23:03

Release: v5.16.0

Release v5.16.0 New Model additions Qwen4-Exp Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE). GR is a Qwen-developed residual architecture that combines Hyper-Connection with GatedNorm. It mixes multiple residual streams with fine-grained elementwise gating before each attention and Mixture-of-Experts (MoE) block, then controls how much of the block output is injected back into each stream. QSA uses multiple query heads to score compressed key blocks, selects the most relevant contiguous token blocks, and keeps the

Transformers 更新✦ 精选规则精选22:50

Release v5.16.1

Release v5.16.1 This is a special release as we include GLM! (and a few small fixes) GLM-5.3-Flash GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply r

8月25日2026-08-25
8月22日2026-08-22
8月21日2026-08-21
Microsoft Research✦ 精选规则精选00:00

Broadening access to Skala creates a faster path to predictive DFT

Skala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to track computational performance. The post Broadening access to Skala creates a faster path to predictive DFT appeared first on Microsoft Research .

8月19日2026-08-19
Transformers 更新✦ 精选规则精选18:50

Patch release: v5.15.1

Patch release v5.15.1 This patch most notably solves a few issues with DFlash and MTP candidate generators, as well as an issue where images could sometimes not be processed on accelerator if using Lanczos filter. It contains the following commits: Fix DFlash candidate token device mismatch with device_map="auto" ( #47877 ) by @sywangyi and @Cyrilvallez Align logit distributions for CandidateGenerators using sampling ( #48007 ) by @Cyrilvallez Fix MTP config when mlp_layer_types is absent ( #48015 ) by @Cyrilvallez Fallback from 'lanczos' to 'bicubic' when on cuda ( #48026 ) by @zucchini-nlp Fix gemma4 video to device ( #47896 ) by @guarin

8月18日2026-08-18
8月17日2026-08-17
8月14日2026-08-14