Share GPU clusters across teams with isolation and fairness using Amazon SageMaker HyperPod

Multiple teams within the same company increasingly need shared access to expensive GPU clusters for their generative AI operations, while maintaining isolation boundaries, resource fairness, and operational independence. Consider a data science team training large language models, a computer vision group running inference workloads, and a research team experimenting with new model architectures. All of … Read more

Introducing Claude Haiku 5.5 on AWS

Today, we’re excited to announce the availability of Claude Haiku 5.5 on Amazon Bedrock and Claude Platform on AWS. According to Anthropic, Claude Haiku 5.5 is the fastest and most efficient model in the Claude 5.5 family, built for subagents and high-volume, cost-sensitive work. It also costs around 75 percent less than Claude Haiku 4.5 … Read more

Rethinking access control for RAG with Amazon Quick and Amazon Bedrock

Enterprise organizations are adopting Retrieval Augmented Generation (RAG) to unlock insights from company knowledge sources like Microsoft SharePoint, Google Drive, and Atlassian Confluence. However, these knowledge sources contain sensitive information governed by complex permission structures. Making sure that AI-generated answers respect those permissions is one of the hardest challenges in enterprise AI. In this post, … Read more

Beyond hours saved: Building the business case for agentic automation

Agentic automation, software that reasons and adapts to complete tasks, is showing up on AI center of excellence (AI CoE) roadmaps. But the standard way companies justify automation investments, hours saved times labor cost minus build cost, was designed for rule-based tools like robotic process automation (RPA). That model misses most of the value agents … Read more

Automate remediation post AWS DevOps Agent investigation

Reducing the time between incident detection, investigation, and remediation is a critical priority for organizations running production workloads on AWS. When an issue arises, on-call engineers often need to quickly diagnose the problem across application components, identify the root cause, and apply the fix, often in the middle of the night. AWS DevOps Agent, an … Read more

Building AI builders: Playbook for closing the AI knowledge-capability gap

The biggest barrier to AI adoption isn’t awareness. It’s the gap between talking about AI and building with it. Professionals whose primary role isn’t writing code, but whose daily work increasingly depends on AI solutions, don’t need engineering backgrounds to start building AI capability. They need the right tools, structured support, and permission to fail. … Read more

How Cornerstone OnDemand cut database diagnosis by 78% with Amazon Bedrock

Cornerstone OnDemand, Inc. (Cornerstone) is a global leader in workforce readiness solutions, serving 140 million users across 186 countries. The company built a multi-agent AI system that transforms database operations from reactive firefighting into proactive, self-orchestrating workflows. The system, called Orion AI, uses Amazon Bedrock and Strands Agents, an open source agent orchestration framework from … Read more

Building a context-aware AI assistant on AgentCore and OpenClaw

Off-the-shelf AI assistants answer individual questions well, but they fall short on a different axis: continuity. Ask a stateless assistant about your garden today and it has no idea that you mentioned your fast-draining raised beds three weeks ago, that you only use organic fertilizer, or that your petunias were struggling through a heat wave. … Read more

Responsible AI governance: How AWS positions customers to align with ISO/IEC 42005:2025

With generative AI adoption moving faster than the personal computer or the internet and global AI-related investment in 2025 representing $581.69 billion, organizations must position their workforce to use AI to power their operations while employing it responsibly. Researchers affiliated with the AI Adoption Initiative argue that the workforce relates to AI through an AI … Read more

Best practices for Amazon SageMaker HyperPod administration and governance

Amazon SageMaker HyperPod gives machine learning (ML) teams access to large pools of accelerated compute for training and fine-tuning models. When several teams share one cluster, the technical setup is usually straightforward. The challenging part is governance. You must decide which teams can use the cluster, how much capacity each team gets, what happens when … Read more

Manage Amazon SageMaker HyperPod Spaces directly from SageMaker Studio

We recently introduced the ability to create and manage Amazon SageMaker Spaces on Amazon SageMaker HyperPod EKS clusters directly from the Amazon SageMaker Studio UI. Data scientists and machine learning (ML) engineers can now launch JupyterLab and Code Editor environments on HyperPod clusters without leaving their browser or using command-line tools, reducing the time from … Read more

Build a voice travel concierge with Amazon Bedrock AgentCore, Managed Knowledge Base and Nova Sonic

Airlines already have apps and websites where travelers check flights, pick seats, and manage bookings, and adding a natural voice layer opens those tasks to spoken requests. With this voice layer, a traveler can change a seat or check a delay by speaking, without leaving the app or navigating through screens. Building it requires careful … Read more

Introducing GLM 5.3 on Amazon Bedrock

Coding and agentic workloads are asking more of AI models than ever: refactor a repository spanning hundreds of files, sustain a multi-hour agentic workflow without losing context, and reason through complex systems problems with tool use at every step. Meeting those demands with open-weight models has historically meant provisioning and operating your own inference infrastructure. … Read more

New agent skill: Amazon SageMaker optimized generative AI inference for your coding agent

Engineers increasingly use coding assistance tools to accelerate their development workflows. Today, Amazon SageMaker AI optimized generative AI inference introduces the aws-ai-ml skill, available through the Agent Toolkit for AWS. This skill gives coding agents like Kiro, Claude Code, and Codex deep expertise in inference optimization and benchmarking. Install the skill, and your existing agent … Read more

Making Amazon Quick enterprise-ready: Automated, auditable cross-account resource promotion

Amazon Quick is Amazon’s agentic AI companion built for work. You build agents that reason over your data, call action connectors, and carry multi-step tasks to completion. Promoting those resources (chat agents, action connectors, knowledge bases, flows, and spaces) from a development to a production AWS account, the way you would any other application, has … Read more

Agentic retrieval with LangChain and Amazon Bedrock Knowledge Bases

When a user asks the support assistant, a Retrieval Augmented Generation (RAG) application built with LangChain to compare two products across three dimensions, they’re effectively posing six questions simultaneously. Similarity search uses a single query vector to encapsulate all the intents. The retriever then generates the best approximation of the average of those intents. The … Read more

Downgrading user roles in Amazon Quick

Managing access permissions effectively is an important aspect of maintaining a secure and collaborative environment in Amazon Quick. Quick supports versatile user management options designed to accommodate various identity types and organizational needs. You can provision users natively through Quick Identity or manage them through enterprise identity providers such as AWS IAM Identity Center or … Read more

Evaluating multi-agent systems for explainability and helpfulness with Amazon Bedrock AgentCore

A critical challenge that emerges as multi-agent systems move from experimentation to production is making sure that these systems are consistently helpful, accurate, and explainable in real-world scenarios. Enterprises are increasingly adopting multi-agent systems to solve complex, real-world problems that require reasoning across data sources, tools, and business constraints. From supply chain planning to financial … Read more

Sweep thousands of leases for compliance using Amazon Quick and the Adjudicated Query pattern

Checking tens of thousands of apartment leases against constantly changing state landlord-tenant laws, and proving you actually checked all of them, has been beyond the reach of most compliance teams. But with generative AI in Amazon Quick, paired with the right backend, it’s now possible. In this post, we introduce a design pattern called Adjudicated … Read more

Add secure Web Search to Claude Desktop with Amazon Bedrock AgentCore

Claude Desktop on Amazon Bedrock provides powerful AI assistance, but without integrated web search, responses are limited to the model’s training knowledge cutoff. When you need current information, such as recent documentation updates, live pricing, or weather updates, the model can’t retrieve it on its own. Amazon Bedrock AgentCore is a platform to build, connect, … Read more

Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI

Search agents powered by large language models (LLMs) are transforming how enterprises retrieve information. Rather than requiring users to craft the perfect query, a search agent autonomously decides what to search for, which retrieval strategy to use, and when to stop searching. It does this across multiple rounds of interaction, refining its approach based on … Read more

Serve live, governed data in AI-built apps with Amazon Quick

Quick Apps could already bring live data into an app from connectors and content sources: action connectors (services like Jira, Slack, and Google Drive), Spaces documents, web search, and AI inference all run at view time, not build time. The solution introduces live structured data from your data lakes, databases, and other analytics data stores. … Read more

Build agent memory with NVIDIA NeMo Agent Toolkit and Amazon S3 Vectors

In my previous post: Building persistent memory for multi-agent AI systems with Amazon S3 Vectors, we explored why memory engineering is the foundational discipline for production multi-agent systems. We showed how Amazon S3 Vectors, a capability of Amazon Simple Storage Service (Amazon S3), meets the architectural requirements for agent memory: semantic retrieval, rich metadata, strong … Read more

Uplifting conversion across the acquisition funnel with personalization using contextual bandits on AWS

Generative AI has made it possible to produce large amounts of personalized content quickly and at low cost. In our previous post, we showed how generative AI on Amazon Bedrock can produce personalized content at scale while staying within brand guidelines and guardrails. The new challenge is now one of selection. Among all of those … Read more

Building ambient agents with Amazon Bedrock AgentCore: From event-driven signals to human-in-the-loop workflows

Teams that process documents at scale know the routine: files land in storage, someone notices, opens each one, decides what it needs, and routes it for review. Monitoring alerts queue up the same way, waiting for a person to act on them. The hours lost to manual triage are the operational problem ambient agents solve. … Read more

Implementing Multi-Environment Access for Claude Platform on AWS

You need Claude Platform on AWS (CPonAWS) inference from three environments: production workloads on AWS, developer laptops for local iteration, and external services on other cloud providers or on-premises continuous integration and continuous delivery (CI/CD) pipelines. Each environment has different authentication requirements, but all should share a single subscription with workspace-level isolation between production and … Read more

Simplify dashboard drill-down with the Amazon Quick Sight hierarchy filter

Amazon Quick Sight is a fully managed, cloud-native business intelligence (BI) capability for building and publishing interactive dashboards. You can access these dashboards from any device and embed them into your applications, portals, and websites. When building a dashboard, authors add filters so readers can narrow the data to what they want to analyze. But … Read more

How uniopen customized Amazon Nova to their retail moderation policies for production deployment

uniopen is a digital communication and membership platform launched by Taiwan’s Uni-President Enterprises Group, connecting customers to ecommerce, membership benefits, and other retail experiences across web, tablet, and mobile channels. Across those channels, uniopen applies a moderation policy that classifies each interaction along two axes. The first is what behavior occurred (nine categories), and the … Read more