Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

On August 12, 2026, Alibaba’s Qwen team released Qwen3.8-2.4T-A95B. This is the first time a Qwen-Max-class model has been made available as open weights. With 2.4 trillion total parameters (95 billion activated per token), a hybrid linear-plus-full-attention architecture, and native context up to 262K tokens (extensible to 1M), Qwen3.8 targets the most demanding agentic and … Read more

ICYMI: What landed for AI builders in August 2026

A recap of the biggest Amazon Bedrock, AgentCore, and Strands updates from August 2026. At AWS, we have long focused on making foundational technologies accessible and providing the infrastructure needed to put them to work. Amazon Bedrock, used by more than 225,000 active customers, including over 80% of Fortune 100 companies, embodies that commitment by … Read more

How Heurist Finance built an AI-native investment workbench on Amazon Bedrock AgentCore

Heurist uses Amazon Bedrock AgentCore to build AI-powered financial intelligence for retail investors. Its flagship product, Heurist Finance, brings several institutional-style workflows into one chat experience: it gathers market data, reads filings and news, runs deep research, builds and stress-tests portfolios, and monitors positions. Each answer reflects the user’s portfolio and preferences. Heurist’s goal is … Read more

Simplify and support your TorchServe workloads using Ray Serve Deep Learning Containers

TorchServe is no longer actively maintained. The official project notice states there are no planned updates, bug fixes, new features, or security patches, and that vulnerabilities might not be addressed. For teams that run model inference on TorchServe today, this means security patches stop and compatibility updates with newer versions of PyTorch and CUDA stop. … Read more

Automate user-level custom permissions for Amazon Quick

As Amazon Quick environments scale and new AI-powered capabilities expand what users can do, automating user-level custom permissions becomes critical to maintaining the principle of least privilege. To address this, with custom permissions in Quick, you can enforce fine-grained access control by toggling specific features on or off for individual users. For example, with custom … Read more

Take on your most ambitious work with GPT-6 Astra on Amazon Bedrock

GPT-6 Astra from OpenAI brings greater depth and judgment to your most demanding tasks and runs on the Amazon Bedrock inference engine built for high performance, security, and scale. Organizations are already running AI agents that write code, analyze data, and automate complex workflows at production scale on Amazon Bedrock. GPT-6 Astra raises the potential … Read more

Pathway’s brain-inspired architecture development on Amazon SageMaker HyperPod

As AI systems take on more complex tasks, much of the industry’s progress has come from increasing model scale, training data, context length, and inference-time computation. Instead of externalizing reasoning work as a chain-of-thought (generating extra tokens sequentially and feeding them back into later steps), Pathway’s brain-inspired BDH (Dragon Hatchling) performs reasoning in latent space. … Read more

Amazon SageMaker Feature Store introduces UpdateRecord for feature-level writes

We are excited to announce feature-level writes for Amazon SageMaker Feature Store. Amazon SageMaker Feature Store is a fully managed, purpose-built repository to store, share, and manage machine learning (ML) features, the processed data used for training models and generating predictions. With the new UpdateRecord API, you can now update one or more feature values … Read more

Govern models with MLflow and Amazon SageMaker AI Model Registry sync: Part 2

Governing models across accounts is the natural next step once automatic model registration is in place. In Part 1 we introduced how managed MLflow on Amazon SageMaker AI synchronizes registered models into the SageMaker AI Model Registry. We walked through a single-account setup where AWS Identity and Access Management (IAM) condition keys separate the data … Read more

Govern models with MLflow and Amazon SageMaker AI Model Registry sync: Part 1

Automating model registration between MLflow and a model registry solves a gap that opens the moment a candidate model leaves experimentation. Data scientists track dozens of candidate runs in MLflow, while governance officers need one authoritative registry to validate, approve, and audit the models that reach production. Managed MLflow on Amazon SageMaker AI already synchronizes … Read more

Automated agent evaluation with Amazon Bedrock AgentCore and GitHub Actions

Build a continuous integration and continuous delivery (CI/CD) quality gate that deploys an agent with role-based MCP tools, evaluates it, and blocks PRs when evaluation scores drop. You shipped an AI agent on Amazon Bedrock AgentCore runtime. It calls tools through an MCP server protected by OAuth. Now you want CI to tell you when … Read more

Benchmarking small LLM inference on SageMaker AI: G7 vs G5 and G6

Choosing the right GPU instance for large language model (LLM) inference is one of the most impactful decisions you make when deploying generative AI at scale. A single generation jump can slash latency, increase throughput, and reduce cost-per-token. However, the real-world magnitude of those gains depends on model architecture, quantization format, and workload shape. In … Read more

How HPE Zerto built an agentic troubleshooting system with Amazon Bedrock

This post was co-written by AWS and the HPE Zerto team. If you manage hybrid and multi-cloud infrastructures, you may already be turning to AI systems to assess health, investigate issues, and act on problems faster. HPE Zerto addressed this challenge by building an agentic troubleshooting system powered by Amazon Bedrock. HPE Zerto Software helps … Read more

How DiDi built intelligent contact center QA with Amazon Bedrock

DiDi partnered with AWS to build an intelligent contact center quality assurance (QA) system on Amazon Bedrock for its International Business Group’s Customer Experience (CX) department. The system covers Spanish and Portuguese across three business lines (ride-hailing, food delivery, and financial services) and migrates QA capabilities from an opaque third-party solution to a transparent, self-owned … Read more

Deploy a multimodal WhatsApp ordering assistant with Amazon Bedrock AgentCore

This post shows how to deploy a multimodal WhatsApp ordering assistant built with Amazon Bedrock AgentCore and Amazon Nova 2. Many quick-service restaurants spread ordering across an app, a website, a phone line, and the counter. Each of those is a separate system to build and run. Each one also fragments the customer’s history, making … Read more

Designing lifecycle policies for AgentCore memory

Memory lifecycle policies help long-running agents on Amazon Bedrock AgentCore stay effective by systematically managing what they remember and forget. Your agent generates memories from every conversation it conducts. If you don’t actively manage these memories, your agents will accumulate outdated context, which can degrade response quality and create compliance risks for your deployment. After … Read more

Build a Physical AI model factory with NVIDIA Cosmos 3 on SageMaker HyperPod

A Physical AI system, such as a robot or autonomous vehicle (AV) that translates real-world data into physical actions, can’t be built in a single training job. Instead, it takes a continuous pipeline: a loop of generating synthetic data, post-training perception and policy models, so the system understands its surroundings and can act, and evaluating … Read more

Run agent-driven Amazon SageMaker HyperPod operations with InstantStart

If you run foundation model (FM) workloads on Amazon SageMaker HyperPod, you know the work is rarely a single task. It is a chain of dependent ones. An infrastructure team creates the network and control plane, attaches accelerator capacity, and installs cluster dependencies in the right order. It also prepares storage and identity, keeps distributed … Read more

Customizing your knowledge base on Amazon Bedrock for large and complex documents using Amazon Textract

For customer service teams handling thousands of utility bills each month, accurately parsing and analyzing complex, multi-page documents is a persistent challenge. Inconsistent formats, dense tables, and varied layouts make it difficult to extract the right information quickly. This leads to delayed responses, billing errors, and frustrated customers. As document volumes grow, these inefficiencies compound, … Read more

How Intuit built an agentic disaster recovery assistant with Amazon Bedrock

Disaster recovery (DR) at scale is hard. When thousands of microservices span multiple AWS Regions, coordinating a reliable failover becomes a major operational challenge. At Intuit, we operate at this scale. We support products that millions of people rely on to run their businesses and manage their finances. These include TurboTax, QuickBooks, Mailchimp, and Credit … Read more

AI-driven development lifecycle using Amazon Bedrock AgentCore

Engineering teams adopting the AI-Driven Development Lifecycle (AI-DLC) with Amazon Bedrock AgentCore and coding agents like Kiro often struggle with the gap between conceptual frameworks and working code. Amazon Bedrock AgentCore is a service for building, connecting, and optimizing agents at scale with any framework or model. AI-DLC positions AI as a central collaborator across … Read more

Set up OpenAI ChatGPT Codex with LiteLLM on Amazon ECS and Amazon Bedrock

OpenAI ChatGPT Codex with LiteLLM can provide centralized enterprise controls for generative AI coding agents. These agents help developers understand repositories, write code, run tests, and complete multi-step engineering tasks. As organizations move from individual experimentation to managed adoption, teams need a consistent way to control model access and attribute consumption. They must also apply … Read more

Best practices for building agentic automations with Amazon Quick Automate

Agentic automations are transforming how enterprises run their business processes. Instead of following rigid scripts, AI agents reason about context, adapt to variation, and collaborate with people and other agents to move work forward. Amazon Quick Automate is a multi-agent automation capability within Amazon Quick that helps organizations build, deploy, and maintain these automations at … Read more

Accessing OpenAI models on Amazon Bedrock from Australia with global cross-Region inference

Australian teams working with OpenAI models can now access the latest OpenAI models through Amazon Bedrock. Amazon Bedrock offers OpenAI GPT-5.6 Sol, Terra, and Luna with global cross-Region inference from both Asia Pacific (Sydney) and Asia Pacific (Melbourne) AWS Regions in Australia. Your application calls the Amazon Bedrock Runtime endpoint in Asia Pacific (Sydney) or … Read more

Modernizing and scaling support operations with generative AI on AWS

Scaling support operations requires handling rising ticket volumes, meeting strict Service Level Agreements (SLAs), adapting to evolving compliance requirements, and maintaining documentation that quickly becomes outdated, all without proportional increases in headcount. In many teams, the knowledge required to resolve tickets is fragmented across SOPs, recordings, and tribal expertise, forcing analysts to spend significant time … Read more

How an AWS team detects dashboard content failures at scale using Amazon Bedrock

Picture a scenario familiar to any organization running business intelligence (BI) at scale: A user opens a dashboard minutes before an important meeting and finds a blank chart. Every infrastructure monitor reports healthy. Servers are up, APIs respond, and the data pipeline completed on schedule. Yet the content on screen is broken, and no monitoring … Read more

From code to diagrams: Agentic architecture documentation with Amazon Bedrock AgentCore

Architecture documentation remains one of the most persistent challenges in software development as code bases evolve rapidly. Development teams often spend hours manually creating architecture diagrams, only to watch them become outdated within weeks of deployment. This documentation gap creates knowledge silos, slows developer onboarding, and complicates compliance audits. Amazon Bedrock AgentCore is the platform … Read more

Trinity: Agentic AI-powered transition planning for students with disabilities

This post was co-authored with Marc Steren, Odina Salihbaeva, and Aashrit Surapaneni from University Startups, a partnership between University Startups, g/d/n/a, and AWS. Trinity is a conversational AI solution that helps students with disabilities take ownership of their postsecondary planning. It was developed by University Startups, which was founded in 2020 on a straightforward belief: … Read more