Spreading the load: How Salesforce met Multi-AZ HA with SageMaker Inference Components

When Salesforce set out to make Agentforce (Salesforce’s AI foundation for agents) highly available (HA) across multiple Availability Zones (AZs), the team faced a gap. Amazon SageMaker AI Inference Components (ICs) could cut GPU costs, but their default placement didn’t guarantee the Multi-AZ resilience Salesforce’s compliance bar required. For Salesforce, the ICs delivered an 8x … Read more

Build agentic creative workflows with Amazon Quick and fal

Creative teams face growing demand for more assets, formats, and revisions, while their scripts, references, models, and outputs often remain fragmented across tools. Creators must repeatedly transfer context and assemble results manually. With 78% of creative leaders saying demand exceeds their teams’ capacity, faster generation alone does not solve the underlying workflow problem. To address … Read more

Introducing India cross-Region inference for OpenAI GPT-5.6 models on Amazon Bedrock

Amazon Bedrock now supports the OpenAI GPT-5.6 models, Terra and Luna, in India, with India geographic cross-Region inference. If you have local data processing requirements in India, including in financial services, healthcare, and the public sector, you can now use these OpenAI models at scale. Amazon Bedrock processes inference requests and data within India. Both … Read more

Deepgram deepens Amazon SageMaker AI observability with Enhanced Metrics

Self-hosted speech AI has historically carried an observability trade-off. The service can tell you an endpoint is up and how many requests it served. The questions that actually drive capacity planning and cost management stay locked inside the vendor’s container: what you are billed for, which features your traffic uses, and what the inference engine … Read more

Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2

This post is a collaboration between AWS, NVIDIA and Heidi. Reducing automatic speech recognition (ASR) inference costs on Amazon Elastic Compute Cloud (Amazon EC2) becomes critical when GPU utilization per request is low but latency requirements are strict. A single ASR inference request typically uses only 15–20 percent of a GPU’s compute capacity, yet the default … Read more

Evaluate any agent framework with Amazon Bedrock AgentCore Evaluations

AI teams building production agents face a frustrating asymmetry: the diversity of agent frameworks keeps growing, but evaluation tooling has not kept pace. Most evaluation systems assume you built your agent in a specific way: a specific SDK, a specific large language model (LLM) client, a specific tracing pattern. The moment you step outside that … Read more

How GoDaddy transformed its analytics with Amazon Quick

GoDaddy is one of the world’s largest domain registrar and web hosting companies, serving more than 20 million customers and managing approximately 82 million domain names. At that scale, access to timely business data directly affects how quickly the company can act. When GoDaddy’s analytics infrastructure faced challenges under the weight of thousands of dashboards, … Read more

Natera’s intelligent appointment scheduling with Amazon Bedrock AgentCore

Booking a phlebotomy appointment shouldn’t be a hassle for oncology patients already managing treatment. Natera’s service, powered by Amazon Bedrock AgentCore, allows a phlebotomist to come to the patient, helping Natera deliver a more convenient experience. Natera, a global diagnostics company specializing in cell-free DNA testing, wanted to transform their patient experience by replacing manual … Read more

Preparing data for supervised fine-tuning Part 2: Advanced data strategies

Data preparation for supervised fine-tuning (SFT) doesn’t end when your dataset is clean and correctly formatted. The harder questions come next. How much data do you actually need? Should you collect more, or select a better subset of what you have? How do you generate high-quality examples when human annotation doesn’t scale? And how do … Read more

Preparing data for supervised fine-tuning Part 1: Formatting and quality

Data preparation determines the ceiling of any supervised fine-tuning (SFT) project. You’ve evaluated your foundation model (FM), and out-of-the-box performance isn’t meeting your production requirements. Maybe the model doesn’t follow your output schema reliably, struggles with your domain’s classification taxonomy, or can’t maintain the tone your application demands. The question isn’t whether to customize, it’s … Read more

Connect Amazon Bedrock AgentCore to cross-account knowledge bases

Organizations often deploy agents using Amazon Bedrock AgentCore, a platform to build, connect, and optimize agents at scale, with any framework or model. These agents may access governed knowledge bases hosted in separate AWS accounts. This cross-account separation helps maintain clear workload boundaries but can introduce integration challenges. This post explains how AgentCore agents in … Read more

Agentic observability with Amazon OpenSearch Service MCP Apps

Observability agents are fast. They query alerts, correlate logs with traces, and produce a root cause hypothesis in minutes. The part that still takes time is verification. You read the agent’s text summary, open your observability tools in a browser, navigate to the trace waterfall, check the service map to scope impact, and cross-reference what … Read more

Governed reports with Amazon Quick Desktop and Amazon FSx for NetApp ONTAP

Amazon Quick Desktop brings governed, AI-assisted reporting to the files your team already manages on Amazon FSx for NetApp ONTAP (FSx for ONTAP), cutting weekly report preparation from hours to minutes. Today, producing those reports takes hours of manual effort each week. Teams re-read the same documents, reformat metrics, and copy summaries into Slack. Leaders … Read more

Introducing new Ray capabilities on SageMaker HyperPod

Today, we are announcing new Ray capabilities on Amazon SageMaker HyperPod that integrate Ray with the HyperPod purpose-built infrastructure for foundation model training and serving. Ray is an open-source framework that data scientists use to scale distributed Python workloads across clusters of GPUs, from distributed training with Ray Train to model serving with Ray Serve. … Read more

Democratizing institutional knowledge: Building an AI-powered knowledge management system with AWS

Organizations across industries struggle with managing institutional knowledge, the collective wisdom and experience accumulated over years of operations. This “tribal knowledge” often disappears when key personnel leave, creating knowledge gaps that impact efficiency and innovation. Traditional documentation methods have proven inadequate, often resulting in outdated or inaccessible information when it’s needed most. In this post, … Read more

Agentic Resource Discovery (ARD): An open specification for agent discovery

How AWS Agent Registry and the Agentic Resource Discovery (ARD) specification enable cross-environment discovery for your agents As organizations scale their use of artificial intelligence (AI) agents and tools, finding the right resource becomes the hard part. Teams build Model Context Protocol (MCP) servers, deploy agents, and create specialized tools, but without a central catalog, … Read more

AI-powered metadata correction and harmonization

As data collection and data generation accelerate, the gap between our ability to produce raw data and our capacity to standardize it continues to widen. Without automation, this gap becomes a critical bottleneck that delays analysis, complicates interpretation, and limits the global value of shared datasets. Metadata harmonization (standardizing labels, identifiers, and formats so datasets … Read more

Agentic Data Operations Platform (ADOP): Data engineering into hours

Data engineering teams routinely spend weeks standing up a single new data source: writing ETL, hand-writing quality checks, updating semantic models, and validating compliance. The Agentic Data Operations Platform (ADOP) on AWS is designed to significantly accelerate that timeline. It’s a reference architecture built on Amazon Bedrock and your AI coding tool of choice. Specialized … Read more

Govern AI agent tool access with Amazon Bedrock AgentCore Gateway

In our conversations with customers over the past months, one pattern keeps recurring. Whether they work with coding agents, autonomous agents, or human-interactive ones, and regardless of workload maturity, we start with the same question: “Which AI agents have access to customer data, who granted it, and what would exposure look like if a credential … Read more

Accelerating aircraft IFEC diagnostics with agentic AI on AWS

Panasonic Avionics Corporation provides in-flight entertainment and connectivity (IFEC) systems across a large global fleet serving hundreds of airlines and billions of passengers annually. When a system issue affects passenger experience at this scale, engineers must diagnose the root cause quickly across thousands of unique deployment configurations. Doing this manually, correlating logs, metrics, and ticketing … Read more

Introducing cross-Region inference for OpenAI GPT-5.6 models on Amazon Bedrock

This post is co-written with Chris Dickens from OpenAI. Amazon Bedrock now offers OpenAI GPT-5.6 models on Amazon Bedrock in more than 25 AWS Regions, with cross-Region inference. Three GPT-5.6 variants support cross-Region inference, Sol, Terra, and Luna, each tuned for a different balance of capability and cost. Cross-Region inference (CRIS) in Amazon Bedrock works … Read more

Build a no-code ML workflow with Snowflake, Amazon SageMaker Canvas and Amazon Quick – Part 1: Setting up your Snowflake environment

Healthcare, retail, and life sciences organizations generate massive quantities of operational data in cloud data warehouses like Snowflake. While these systems store and scale information efficiently, transforming that data into meaningful predictions remains a challenge. Traditional machine learning (ML) approaches require specialized teams, long development cycles, and heavy engineering support, creating delays and limiting experimentation … Read more

Build a no-code ML workflow with Snowflake, Amazon SageMaker Canvas and Amazon Quick – Part 2: Data preparation and model building with Amazon SageMaker Canvas

Part 1 covered the Snowflake database setup and established the foundational infrastructure for this no-code machine learning (ML) workflow. Part 2 of this blog series covers complete data preparation and model building workflow using Amazon SageMaker Canvas, demonstrating how to connect directly to Snowflake data sources, transform and prepare data using Data Wrangler’s visual transformations, … Read more

Build a no-code ML workflow with Snowflake, Amazon SageMaker Canvas and Amazon Quick – Part 3: Visualizing insights with Amazon Quick Sight

Part 1 covered the Snowflake database implementation setup and established the foundational infrastructure for our no-code machine learning (ML) workflow. Part 2 walked through the complete data preparation and model building workflow using Amazon SageMaker Canvas, demonstrating how to connect directly to Snowflake data sources, transform and prepare data using Data Wrangler visual transformations, and … Read more

Authoring Dogwood policies from natural language in Amazon Bedrock AgentCore

AI agents can automate complex workflows but might take actions that don’t align with your organization’s policies or regulatory constraints if used without proper controls. To address this, we built Policy in Amazon Bedrock AgentCore so teams can implement controls that are applied across agents running in Amazon Bedrock AgentCore. This was recently expanded with … Read more

Scaling agentic AI: Enterprise patterns without vendor lock-in

Scaling agentic AI across an enterprise requires architectural patterns that preserve flexibility while avoiding vendor lock-in. This post is Part 2 of our series on multi-agent systems at scale. In this post, we examine how machine learning (ML) teams operate agentic AI systems across a “multi-everything” environment of frameworks, models, and providers. We also cover … Read more

Scaling cloud migrations with agentic AI on Amazon Bedrock AgentCore

Scaling cloud migrations with agentic AI on Amazon Bedrock AgentCore starts with recognizing where large-scale migrations break down. Discovery consumes weeks per application. Engineers write infrastructure code from scratch for each workload. Post-migration operations devolve into reactive firefighting. Multiply those bottlenecks across over 300 applications and a fixed fiscal year deadline, and migration programs struggle … Read more