Scaling MoE reinforcement learning on Amazon EKS with EFA and DeepEP with 40% more throughput

When you post-train a Mixture-of-Experts (MoE) model with Reinforcement Learning from Human Feedback (RLHF) or Group Relative Policy Optimization (GRPO) at scale, three simultaneous challenges emerge. The first requires coordinating heterogeneous compute for rollout generation and policy training. Second, sustaining high-throughput communication across hundreds of accelerators. And third, dynamically orchestrating every subsystem to keep them … Read more

Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod

Reinforcement learning (RL) post-training is becoming a standard step in building capable language model agents. Models learn to reason and act across sequences of steps by generating trajectories, receiving rewards, and updating their policy based on outcomes. Running this at scale, across multiple nodes with hundreds of GPU-hours of rollouts per training run, requires persistent … Read more

NarrateAI: production-ready LLM quality assurance on Amazon Bedrock

Executives need to make data-driven decisions during live business reviews, where accuracy and speed matter. A conversational agentic AI assistant can meet this need by answering data questions instantly. But the stakes are high: a wrong number or a slow response in front of leadership carries immediate professional consequences, and a capable large language model … Read more

Deploying real-time personalized speech with Qwen3-TTS on Amazon SageMaker AI

With voice cloning, you can generate new speech in a target speaker’s voice from a short reference recording, without retraining a model. You can now deploy the publicly available Qwen3-TTS-12Hz-1.7B-Base text-to-speech model from Amazon SageMaker JumpStart to a fully managed, real-time inference endpoint. Voice cloning reproduces the vocal identity of a specific speaker. Start with … Read more

How Datacor built self-service rental analytics with Amazon Quick Sight

This post was written with contributions from Datacor’s TrackAbout engineering and product teams. For gas and welding distributors, rental billing on assets such as cylinders and bulk tanks is a significant share of total revenue. Yet the data needed to manage those assets was often locked in disconnected systems, accessible only through IT. The resulting … Read more

Multi-Region training with Amazon SageMaker HyperPod and Qumulo

With Amazon SageMaker HyperPod and Qumulo, you can place training compute in one AWS Region and keep your dataset in another. Training large AI models requires massive GPU capacity, but your ideal compute resources and your training data don’t always reside in the same AWS Region. Accessing data across Regions adds network latency and transfer … Read more

Speaker-labeled transcription with WhisperX on SageMaker AI

Any team working with spoken audio hits the same wall with generic speech-to-text. Think contact-center calls, all-hands meetings, podcasts, depositions, and broadcast media. These workloads need two things that standard transcription gets wrong. First, timestamps land at the utterance level, off by several seconds. Second, there’s no reliable answer to “who said what.” Those gaps … Read more

Build a multi-account AI agent with AgentCore Gateway and MCP

Enterprises increasingly want AI agents that can reason over data spread across many AWS accounts without copying or centralizing it. Each team keeps its data in its own account for good reasons: clear ownership, scope isolation, and independent deployment lifecycles. But an agent that sees only one account’s data delivers limited value, and connecting it … Read more

Aderant builds intelligent ticket triage with Amazon Nova

This guest post is co-written by Angela Mapes and Adam Walker of Aderant. In this post, we share how Aderant, a global provider of business management software for the legal industry, built an intelligent ticket triage system using Amazon Nova Lite through Amazon Bedrock. Aderant’s solution automates much of the context gathering, classification, routing, and … Read more

From portal-hopping to instant answers: HEMA’s journey with MCP and Amazon Bedrock

This post is co-written with Mauro Rallo and Patrick van der Plas from HEMA. When engineers at HEMA needed an answer, they went portal-hopping, navigating disconnected wikis, service catalogs, and IT portals to find it. To turn that friction into instant answers, the 100-year-old Dutch retailer built a knowledge layer on Amazon Bedrock AgentCore. HEMA … Read more

Agentic conversational video intelligence built on AWS

With video intelligence powered by agentic AI, you can ask natural language questions about uploaded videos and get answers within seconds. Organizations across media, security, insurance, and professional services are generating more video than their teams can review. Meeting recordings accumulate in shared drives, and security cameras capture weeks of unreviewed footage. Field inspection videos … Read more

Claude Opus 5.5 is now available on AWS

Today, we’re excited to announce the availability of Claude Opus 5.5 on Amazon Bedrock and Claude Platform on AWS, the first of the Claude 5.5 model family. Claude Opus 5.5 is Anthropic’s most capable Opus model suitable for agentic coding, knowledge work, and long-running tasks. This post covers Claude Opus 5.5’s improvements, practical guidance, and … Read more

Evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore

General-purpose agents handle a broad range of tasks, but you still need them to follow the procedures that run your business: compliance checks, document-processing workflows, escalation policies, engineering conventions. Encoding all of that in one system prompt or in application logic gets hard to maintain and update. Skills are a modular alternative. A skill is … Read more

How Reactiv automates mobile commerce 80% faster with Amazon Bedrock AgentCore

Reactiv offers a mobile commerce product that helps Shopify merchants launch and manage native mobile apps, where shoppers convert at 2–4 times the rate of web visitors. For these merchants, a stale homepage or a missed promotional window costs real revenue. Yet keeping an app fresh requires constant manual work: choosing which products to feature, … Read more

Right-size generative AI endpoints with concurrency sweeps on Amazon SageMaker AI

Concurrency sweeps help you right-size a generative AI endpoint by finding the instance type and serving configuration that maximizes price-performance while holding latency within acceptable bounds. Without a systematic approach, right-sizing means deploying, load-testing manually, adjusting, and repeating until the numbers look acceptable. Choose five ml.g7e.2xlarge instances when one would suffice, and you burn your … Read more

How Trane gets building insights 60x faster with Amazon Bedrock AgentCore

Trane Technologies manages millions of connected heating, ventilation, and air conditioning (HVAC) assets worldwide, but getting a single operational answer could mean cross-referencing multiple dashboards and drilling through menus for 20 minutes or more. For organizations operating at this scale, that kind of friction slows operations, defers corrective action, and creates material business impact across … Read more

Extending public sector intelligence with Agentforce and AWS

Public sector agencies process large volumes of unstructured evidence, such as body camera footage, surveillance video, and scanned documents, that require extracting insights before anyone can act on them. This post shows how to combine Amazon Bedrock Data Automation with the Model Context Protocol (MCP) to turn unstructured data into structured insights. You can then … Read more

How BMW Group detects cost anomalies across 14,000 cloud accounts

This post is co-written with Philipp Karg from BMW Group and Christopher Masurek from Data Reply. Cost anomalies are hard to spot when you run 14,000 cloud accounts. BMW Group operates Cloud Efficiency Analytics (CLEA), an in-house FinOps system built on AWS with Reply that monitors more than 14,000 cloud accounts across BMW Group’s cloud … Read more

Run Positron on Amazon SageMaker AI for data science workflows

Data science teams often move among separate tools for governed data access, R analysis, Python model development, deployment, application development, and reporting. Positron, Posit’s integrated development environment (IDE) for data science, now runs on Amazon SageMaker AI. For a data scientist, running Positron on SageMaker AI means: Data access without managing credentials. Positron runs under … Read more

How Benchling secured multi-tenant AI agents with Amazon Bedrock AgentCore

When Benchling needed to run AI agent-generated scientific code across thousands of life sciences tenants, their security team found that traditional sandboxing wasn’t enough. Today, this architecture processes more than 600 code execution sessions per day across more than 250 tenants per week with zero security incidents. Standard network controls block HTTP, restrict egress ports, … Read more

Reducing medical claims review time with AI on AWS: The EXL Medical IDP solution

Insurance claims adjusters spend over 100 minutes per case manually reviewing medical records. The EXL AI-powered Medical intelligent document processing (IDP) solution, built on AWS, transforms this process. It combines IDP with domain-specific large language models (LLMs) to extract, summarize, and query medical information at enterprise scale. Challenge: Medical records are complex, voluminous, and critical … Read more

Amazon SageMaker Inference: 2026 year-to-date launches in review

Generative AI inference is uniquely hard: models are tens to hundreds of gigabytes, latency requirements are measured in tokens per second, cold starts can span multiple minutes as containers and weights transfer, GPU capacity is constrained, and traditional monitoring tools expose none of the token-level signals that matter in production. Amazon SageMaker AI offers customers … Read more

Introducing Kimi K3 on Amazon Bedrock

Open-weight models are changing the economics of building and deploying AI at scale. Rapid gains in intelligence and efficiency mean companies can match each workload with the right balance of capability, speed, and cost. AWS is building for a future in which organizations can adopt open-weight innovation with the reliability and security required for production. … Read more

Migrating multi-model AI agents to Amazon Bedrock AgentCore runtime

Organizations building multi-model agentic AI applications face growing infrastructure complexity. Managing container orchestration, scaling policies, identity, and observability for multiple model types adds operational overhead. Teams often spend more time on infrastructure than on agent logic development. Developers running agentic frameworks on self-managed infrastructure such as Amazon Elastic Container Service (Amazon ECS) with AWS Fargate … Read more

The new AgentCore runtime: Elastic, optimized, and consistently fast starts

Agents are no longer experiments. They process claims, write and review code, coordinate across systems, and run for hours without supervision. As agents take on more complex, longer-running work, the infrastructure underneath them must evolve just as fast. We built Amazon Bedrock AgentCore to help developers build, connect, and optimize agents securely at scale. AgentCore … Read more

Deploy Hugging Face models on Amazon SageMaker AI with coding agents

Deploying a Hugging Face model to production means making a dozen decisions: choosing the right serving container for the model’s architecture, confirming the current image tag for your AWS Region, and matching an instance type to the model’s memory footprint. Beyond infrastructure, you must wire autoscaling so you don’t burn GPU hours on an idle … Read more