全部 AI 动态
Runtime instances: persistent compute for production AI agents on Amazon Bedrock AgentCore
Announcing runtime instances in Amazon Bedrock AgentCore—persistent, managed EC2 infrastructure for production AI agents with multi-agent collaboration, GPU support, and sessions lasting up to 14 days.
Resilient HPC and ML on AWS: Running Tightly Coupled Workloads on Spot Instances
This post references AWS ParallelCluster. Check out AWS Parallel Computing Service (AWS PCS), our new managed Slurm service for running HPC and AI workloads on AWS. This post was contributed by Santosh Kumar, Bhagyaraju Kasina, Dr. Sandeep Sovani and Dr. Max Starr Researchers and engineering teams running High Performance Computing (HPC) jobs face a constant […]
Towards a quantum computer that learns from its errors
Machine Intelligence
We’re announcing the Alliance for America’s Skilled Trades.
Google is joining BlackRock, Carhartt and Ford to launch the Alliance for America’s Skilled Trades.
Patch release: v5.14.1
Patch release v5.14.1 This patch solves a few issues which appeared when integrating Inkling model, most notably an issue affecting models using EncoderDecoderCache during assisted generation. It also fixes an issue that could appear during prefill with StaticCache and sdpa without padding for Inkling which uses a position_bias. It contains the following commits: Fix sdpa prefill with position_bias ( #47359 ) by @Cyrilvallez Fix assisted decoding for models with EncoderDecoder cache & OlmoHybrid ( #47361 ) by @Cyrilvallez [FP8] Bump kernels version ( #47344 ) by @vasqu Fix deepgemm on multiple devices ( #47323 ) by @IlyasMoutawwakil
Building a Custom Metrics Exporter for Kubernetes
Kubernetes ships with built-in awareness of CPU and memory, but most real-world scaling decisions depend on signals that live entirely outside that narrow window: how many messages are waiting in a queue, how long the last batch job took, how many active WebSocket connections a pod is holding. When the built-in metrics are not enough, a metrics exporter bridges that gap. This post walks through writing one from scratch, packaging it as a container, and wiring it into a cluster so that Prometheus — and ultimately the HorizontalPodAutoscaler — can consume it. What a metrics exporter actually does An exporter is a small HTTP server with a single
Our largest solar and battery storage project ever
A group of people breaking ground on a construction site
Patch release v5.13.1
Patch release v5.13.1 This patch is focused on enabling transformers for the latest release of vllm! Be more defensive with remap_legacy_layer_types for custom models ( #47245 ) from @hmellor Fix custom code which doesn't know about the new linear layer type names ( #47174 ) from @hmellor Fix case where _LazyAutoMapping.register is passed a str key ( #47148 ) from @hmellor
Transforming HPC Operations with Intelligent Workload Orchestration on AWS
This post was contributed by Manu Pillai, Gloria Macia and Natalia Jimenez, PhD Organizations running high-performance computing (HPC) workloads today operate largely as they have for decades: users manually specify required compute specifications for each of their jobs. Users spend valuable time analyzing workload requirements, selecting instance types, and troubleshooting infrastructure issues – time that […]
Optimizing cloud economics with linear elastic caching
Algorithms & Theory
Spotlight on WG Device Management
The rising popularity of AI, Edge, and Telecommunications workloads on Kubernetes has led to new requirements for hardware management. We now need hardware specification beyond CPU time and memory allocations. This includes allocating GPUs, TPUs, network interfaces, and other hardware, sometimes after pod start and occasionally through time-sharing. Efficiently managing this specialized hardware is the mission of the Device Management Working Group . Their cornerstone project, Dynamic Resource Allocation (DRA) , recently graduated to GA, marking a fundamental shift in how the project handles hardware-intensive workloads at scale. In this spot
Patch release v5.12.1
Patch release v5.12.1 Updated the lower bound for PEFT and a fix for auto tokenizer to properly resolve the mistral tokenizer (when mistral-common is installed). This is similar to v.5.10.3 minus the fixes that were already included in the main release - vLLM will first target 5.10.3 🤗 Fix peft lower bound #46605 by @hmellor ( #46605 ) mistral common backend fix #46667 by @itazap ( #46667 ) Full Changelog : v5.12.0...v5.12.1
A low-carbon computing platform from your retired phones
Climate & Sustainability
Our new community investments in Virginia support local jobs and expand energy affordability.
We’re helping build the state’s next-generation workforce and investing in energy programs.
Growing the next generation of American workers
A man poses in a hard hat
Reducing costs by 50% while processing population-scale genomics with Mountpoint for Amazon S3 and AWS Batch
This post was contributed by Kambiz Shahim, Ankit Kalyani, and Chris Wright. Oxford Nanopore Technologies used Mountpoint for Amazon S3, AWS Batch, and Nextflow to build EPI2ME Cloud, a managed compute environment for processing human genomes at population-scale reliably and securely while reducing computational costs by 50%. EPI2ME Cloud forms Oxford Nanopore’s suite of local […]
We’re announcing a new data center and energy investments in Gray and Roberts Counties, Texas.
Google and Intersect are announcing construction of the Meitner Energy Center, a new data center and new energy generation in Texas.
Google’s water stewardship commitments for local communities
Sunrise at Brenton Slough, showing a body of water surrounded by green land and trees
Monitoring AWS Parallel Computing Service
This post was contributed by Ronald Hudson and Nate Haynes High Performance Computing (HPC) on AWS demands precise monitoring, like the racing telemetry used by Formula 1 teams to deliver results. Like race engineers tracking car performance, AWS Parallel Computing Service (AWS PCS) administrators must monitor computing metrics in real-time. This vigilance is critical because […]
Blue, yellow and green: Google invests in its new data center in Sweden.
A final render of the data center.