System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics. Higher system performance means more tokens generated, resulting in higher revenue. Efficient scaling means throughput grows proportionally as hardware gets added, requiring fewer resources to serve users at scale. Continuous optimization means generating more value from infrastructure investments. […]
Know everything. Do anything. That was the message NVIDIA founder and CEO Jensen Huang brought to Salesforce Dreamforce Tuesday, joining CEO Marc Benioff onstage in an appearance that coincided with the announcement of Koa — Salesforce’s first CRM reasoning model, built on NVIDIA Nemotron 3 Super. Huang didn’t just take the stage. He walked into […]
Power is a defining constraint for AI factories. As AI workloads demand a full compute platform to serve them, each component of that platform must maximize...
Biomolecular structure prediction is now often run at proteome scale, where the goal is to move an entire worklist through the pipeline efficiently. NVIDIA...
Encode-prefill-decode (EPD) disaggregation is an inference optimization technique for multimodal models that separates the vision encoder stage from the prefill...
OpenAI spent Sunday talking about recursive self-improvement. Chief Scientist Jakub Pachocki's essay An Alien Mind says internal results give him a strong expectation that progress can be sustained into RSI, that chain-of-thought monitoring is becoming progressively less reliable on the Astra class, and that OpenAI will unilaterally withhold scaling if needed while calling for mandated safety bars. The companion data post is the more concrete document: the median OpenAI researcher now burns over $600 a day of inference at API prices, the 90th percentile over $7,000, the research org runs 3.1 agent-workdays per human workday, and a July 20 inf
The first full weekend with GPT-6 Astra in everyone's hands, and the interesting findings are the unglamorous ones nobody put in a launch video. You can now change reasoning effort mid-conversation without invalidating the prompt cache , because the effort change is appended to the end of the context instead of rewriting the top. Codex has an experimental compaction mode where Astra keeps notes across context windows and can search earlier windows including tool calls, off by default and buried in a TOML flag. The hallucination-rate drop is on pages 20 and 21 of the system card and OpenAI barely mentioned it. Third parties are filling in the
Running reasoning and agentic AI at the edge has been harder than it needs to be. Until recently, models capable of multi-step reasoning were too large to run...
OpenAI shipped GPT-6 Astra , priced exactly like Fable at $10 in and $50 out, and the day split three ways. The capability story is real: 99.9% on ARC-AGI-3, two Lean-verified Erdős problems no model had touched, a prime-gap bound improved for the first time since the 1930s, and Latent Space's writeup after 20 billion tokens calling it an AI engineer you can hire for under six dollars an hour. The benchmark story is messier: the ARC score needs OpenAI's own harness that preserves hidden reasoning state (the standard harness gets 62.7%), Artificial Analysis has Astra level with GPT-5.6 Sol and five points behind Fable 5.1 on general intelligen
AI agents are learning to do more by working together. A lead agent can break a complex task into smaller jobs and assign those jobs to specialized subagents....
This post is the third in a series on AI model co-design. It explores how to accelerate LLM inference while maintaining accuracy using speculative decoding and...
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
The surge in AI adoption is transforming everything from chatbots to content generation. Still, a common pain point remains: How can organizations confidently...
IsoExec unifies numerical execution across SkyRL's vLLM and Megatron runtimes, reducing the average rollout-versus-training logprob difference below 1e-6 on Qwe
A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research .
Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research .