精选
Seeing beyond BMI: Estimating cardiometabolic risk with smartphone imagery
General Science
Introducing Gemini 3.7 Flash
What We Learned by Reproducing 2,200 papers from ICML
Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis
MindTopo reveals VLMs’ spatial reasoning abilities
A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research .
Empty shelves or lost keys? Recall is the bottleneck for parametric factuality
Generative AI
Advancing AMIE towards expert-level audio-visual clinical consultations
Health & Bioscience
Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research .
Thinking of ACE? We Can Do It with Fewer Tokens
Release: v5.15.0
Release v5.15.0 New Model additions Meta Muse Glimmer Muse Glimmer, released today, is Meta’s new multimodal model, especially designed for agentic use cases. Distilled from Muse to 30B parameters, and released under the Apache 2.0 license, it can be deployed to local setups for privacy-aware applications such as coding, document analysis, personal assistants, Claw- or Hermes-like setups. Muse Glimmer is a dense 30B parameter model consisting of: 2B ViT-style encoder for vision (Perception Encoder) 28B parameter text decoder We're covering it in the following blogpost: http://hf.co/blog/muse-glimmer GraniteMoeSWA & GraniteSWA Links: Documenta
Making Knowledge Distillation Cheap Enough to Run at Scale
Meta is back with Muse Glimmer: local, agentic, multimodal, and open source
WeatherNext: AI model achieves breakthrough in forecasting cyclones
EvoLib: Turning experience into evolving knowledge
LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib: Turning experience into evolving knowledge appeared first on Microsoft Research .
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.
Introducing Gemini 3.5 Flash Cyber
Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.
Release v5.14.0
Release v5.14.0 New Model additions Inkling (fresh from Thinking Machines): 975B total, 41B active Add Inkling model #47347 by @molbap @Cyrilvallez @eustlb and @zucchini-nlp Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI- powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-foll
Towards demystifying the creativity of diffusion models
Algorithms & Theory
SensorFM: Towards a general intelligence and interface for wearable health data
Generative AI
Release v5.13.0
Release v5.13.0 New Model additions KimiK 2.5, 2.6, and 2.7 This release includes the architecture for Kimi 2.5 which is used by 2.5-2.7: Kimi K2.5 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration. The model was proposed in Kimi K2.5: Visual Agentic Intelligence and further improved in [Kimi K2.6: Advancing Open-Source Coding](Kimi K2.5: Visual Agentic Intelligence). Kimi K2.5 achieves significant improvements on complex, end-to-end coding tasks, generalizing robustly across programming langua
Introducing TabFM: A zero-shot foundation model for tabular data
Data Management
Thinking to recall: How reasoning unlocks parametric knowledge in LLMs
Generative AI
Research into how AI can help users understand skin conditions
Health & Bioscience
DiffusionGemma: 4x faster text generation
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
Towards passive heart health monitoring via smartphone camera
Health & Bioscience
The next chapter in flood resilience: Open sourcing Google’s hydrology framework
Climate & Sustainability
A New Era of Discovery: Google Research at I/O 2026
General Science