Amazon Web Services (AWS) is registering its fastest expansion in nearly five years, propelled by a sharp increase in artificial intelligence workloads across both high-end research labs and a broad spectrum of enterprise customers, AWS CEO Matt Garman said in an interview on Bloomberg TV's Bloomberg Tech.
Garman described a demand environment unlike anything AWS has encountered recently. He outlined how organizations ranging from frontier AI labs to finance, healthcare, retail and media companies are deploying AI models at scale, prompting AWS to accelerate its investments in custom silicon and physical infrastructure to meet customer needs.
AI growth across a wide customer base
While some of the most publicized activity comes from prominent research labs, Garman emphasized that the increase in AI compute is not concentrated only among those few large customers. Startups and established enterprises alike are embracing AI to improve operations and deliver new services. According to Garman, AWS’s growth is therefore distributed across many customers rather than depending disproportionately on a handful of very large users.
$25 billion AI run rate shifting toward inference
Garman said AWS’s AI business is operating at about a $25 billion revenue run rate. That total reflects both training of models by major AI developers and an expanding volume of inference workloads - the process of running trained models in production to deliver real-world functionality. He noted that customer spending is steadily moving toward inference, which is increasingly the driver of direct business value for end customers and is being delivered through services such as Amazon Bedrock.
Capacity shortages and elevated capital spending
To keep pace with demand, AWS has significantly increased capital expenditure. Garman confirmed that Amazon’s overall capex stands at $220 billion this year - an increase of $20 billion - and that the company will continue heavy capital spending into next year. He said demand materially exceeds available supply, prompting customers to lock in long-term arrangements.
- Many of AWS’s compute commitments are already reserved through the end of 2027 and extend well into 2028.
- Customers are entering five-year commitments to secure the compute capacity they expect to need for AI workloads.
Custom silicon - Trainium and Graviton
Part of the $25 billion run rate derives from capacity AWS rents that is powered by its own processors rather than from direct chip sales. Garman outlined the role of AWS’s in-house silicon efforts in meeting AI demand.
- Trainium capacity is largely sold out through the end of next year, according to Garman.
- By controlling the hardware and software stack, AWS says customers can achieve cost savings on inference of roughly 20% to 30% using Trainium compared with conventional market options.
- AWS continues to offer Nvidia GPUs alongside its own silicon, remaining one of Nvidia’s largest customers, and provides both options so clients can choose the hardware best suited to their workloads.
- Garman indicated AWS currently rents compute capacity rather than selling chips, though he left open the possibility that the company could consider selling chips outright to third parties at some future point.
Open-weight models and regulatory stance
Garman discussed AWS’s support for an open-weights letter and framed the company’s position on regulation as seeking balance. He said oversight should be applied consistently across closed frontier models and open-weight models to avoid hampering innovation. He also suggested that creators of open-weight models will increasingly move toward licensing arrangements when those models are deployed in commercial cloud settings, as developers look to monetize their intellectual property.
Outlook
The interview presents AWS at a pivotal operational juncture: record-setting demand for AI compute is colliding with finite physical capacity. With significant portions of compute capacity committed years in advance and capex elevated to build additional infrastructure and expand custom silicon deployment, AWS is positioning to meet near-term enterprise demand while continuing to manage supply constraints.
How quickly supply can be expanded to align with accelerating demand will shape the company’s operational priorities in the months ahead, including continued heavy capital investment and long-term customer contracts to secure compute resources.