Building the Cloud for an Agentic World | AWS CEO Matt Garman
- 01AWS is retooling the cloud for agents as primary users, not just people
- 02Frictionless onboarding: removing the VPC/IAM/credit-card tax for new accounts
- 03The capex supercycle: $220B and a multi-constraint supply chain
- 04Diversification as a hedge against the "AI bubble"
- 05Startups are the strategic feeder system, with deliberate GPU allocation
- 06Vertical integration into silicon: Nitro lineage to Graviton to Trainium
1. Key Themes
AWS is retooling the cloud for agents as primary users, not just people
Garman argues agents have different needs than human developers: tail latency, instant provisioning, short-lived permissions, and a frictionless onboarding path. AWS is both optimizing existing services and building new primitives.
Garman: "Agentic workflows tend to perform better on AWS than anywhere else." 00:00:00 On tail latency: "They actually care a lot about tail latencies, which is interesting... they don't always care about the P999 S3 latency, like agents do care and get blocked by that." 00:08:20 On new primitives: "compute sandboxes, gateways, agent permissions versus people or service role permissions. A lot of those are things that we have built and are building and thinking actively about because they are just brand new building blocks." 00:00:00 He also highlights "AWS Context" (in beta), a layer that lets agents find data across Aurora, S3, and other stores. 00:07:51
Frictionless onboarding: removing the VPC/IAM/credit-card tax for new accounts
AWS is rolling out account creation that works in under 30 seconds with just a Gmail address, with enterprise controls defaulted behind the scenes and no migration later. This is a direct response to agents and startups choosing the path of least resistance.
Garman: "you don't have to give it a credit card, you can sign up with your Gmail account... within less than 30 seconds, you're up and running and can be operating a full AWS account." 00:10:13 And critically: "it's not like a simplistic account and then you have to migrate... how do we make it not a choice for the customer, but an easy on-ramp into the depth of features." 00:10:43 He also concedes that for "just try it" deploys, agents often route to partners with easier layers on top: "a lot of times that'll go to some of our partners that have an easier to use layer on top, which we love by the way." 00:09:44
The capex supercycle: $220B and a multi-constraint supply chain
Amazon's spend is unprecedented, and the binding constraint keeps rotating among power, memory, chips, construction labor, and capital.
Garman: "$220 billion for 26... that is a larger expense than we've ever had, maybe any company has ever had in a single year. And we don't anticipate slowing down anytime soon because the demand is just massive." 00:00:41 On constraints: "whether it's power, data centers, capital, memory, chips, they're all kind of constraints at various times. Construction people that build buildings... construction people are at a premium today." 00:17:18 Quoting The Goal: "there's never one constraint there's always just the latest constraint." 00:27:03 AWS now plans years out, not quarters: "we work multiple years out... what do we need for 26, what do we need for 28." 00:25:33
Diversification as a hedge against the "AI bubble"
AWS positions its customer breadth as structural protection versus neoclouds with extreme concentration.
Garman: "some of these... NeoClouds or some of the other providers... you'll see sometimes concentrations of 30, 40, 50, 60% with one or two customers. We're nowhere near that obviously, single digit percentages at the highest." 00:21:17 On demand quality: he says customers asked whether they see positive returns "almost to a person they'll say, oh yeah," so "there's no bubble in which they stop spending on that." 00:22:17 He also argues AWS's production-workload gravity matters: "AWS is where people really are coming to launch their production workloads." 00:21:47
Startups are the strategic feeder system, with deliberate GPU allocation
Despite being able to sell every GPU to frontier labs, AWS reserves capacity for startups because they are tomorrow's enterprises and the source of product learning.
Garman: "we estimate that maybe 30, 40% of AWS revenue comes from companies that was a one-time startup in AWS's lifetime." 00:03:57 On allocation: "we could sell every single GPU or AI accelerator we had to probably just the big frontier labs... We choose not to do that because we actually want to keep growing the full ecosystem." 00:18:17 "we say... yes, in some way, shape or form to something like 60% of the requests we eventually get." 00:18:47 He notes startups push the frontier ahead of banks and governments: "the banks are going to want some capability five years down the road that startups want today." 00:04:55
Vertical integration into silicon: Nitro lineage to Graviton to Trainium
AWS's chip strategy grew from offload cards and incrementally proven, low-risk steps, and is now a core competitive weapon in inference.
Garman: virtualization offload gave better performance, utilization, and security isolation: "we tell people we have no access to any of your VMs that are running there and this has been a huge benefit for us for the last decade where frankly we have been leading and others have been slow to do this." 00:34:20 Graviton value: "they're 20% cheaper at a 20% better performance and have been like that for the last five six years." 00:35:52 "We have examples where people have moved their whole fleets and cut the number of servers they had in half." 00:36:23 On Trainium: "we're now in market with third generation Trainium 3... we're sold out for capacity probably towards the end of next year." 00:36:53 And: "it turns out that Trainium is actually maybe the best inference chip on the market right now." 00:37:47
The enterprise agent gap: reimagine workflows, then earn trust to go autonomous
Enterprises are stuck on two things: replicating human workflows instead of redesigning them, and trusting agents with autonomy.
Garman: enterprises said "great, I would have an agent go do the same workflow" and AWS pushes them to think greenfield rather than "Bob does step one two three four five so agent is going to check it." 00:40:39 On trust: "you don't actually want an agent to just go crazy and accidentally delete a production database. That's going to be pretty bad." 00:42:10 On evals: "How do you have a constant loop of testing? How do you think about goal seeking... All of those things are problems that enterprises don't know how to solve today. I don't know if anyone really is great at solving these today." 00:43:12
A teach-to-fish forward-deployed engineering model
AWS explicitly rejects the perpetual-consultant model; FDEs go in for a bounded window and leave the customer self-sufficient.
Garman: "in 45 days do work where we can teach them how to make an eval, teach them how to get their data in a labeled way... at the end of 40 to 5 days you leave and that customer is good and ready to go and trained up." 00:44:11 "They don't want to be beholden to an external workforce for the next five years but they need help today." 00:44:11
Data sovereignty as the Bedrock moat, plus the open-weights-on-SageMaker comeback
Bedrock's pitch is that prompts never reach the model provider, and AWS was criticized early for slowness that turned out to be deliberate.
Garman: "we have a guarantee that your data never leaves your VPC... they never see your prompts." 00:45:10 "three years ago I got a lot of heat... for being slow to the AI world because we actually built the foundations... now as people move from proof of concepts to production vast majority of them are landing in AWS on Bedrock." 00:45:40 On open weights: customers believe "if they could mix in, do some post training, do some fine tuning to an open weights model that they could distill down they can actually get a better performing model at a lower price," and "most people are doing that in SageMaker today." 00:47:41
AWS runs on agents internally: product velocity and smaller, mobile pods
Agent-first development, not code completion, is changing AWS's own shipping cadence and org structure.
Garman: "it's not code completion it really is agent first, the agents write all of the code, you're just managing a team of agents." 00:53:19 On org design: "one doesn't have to be 10 people it can be three to four people and they build something so fast you actually want to move them to different projects." 00:54:10 Non-engineers benefit too: HR agents turn "what used to take teams of people weeks" into "a single person... in a couple of hours." 00:51:42
2. Contrarian Perspectives
Durability can be over-engineered for agents, yet AWS refuses to ship a "non-durable" option
Most would assume hyperscaler-grade durability is always a feature. Garman admits that for ephemeral agent databases, five nines of durability is arguably wasted. But his answer is not a cheap, lossy tier; it is making creation instant and disposal cheap while retaining durability, because you can't know which databases will become production.
Garman: "many agents want to create a database, do a little bit of work, and then have the database go away. You really need that database to have five nines of durability. Exactly. Right?" 00:12:26 Then: "we don't really want to have like a non-durable option that's going to cause problems easier because you never really quite know if the database that's created wants to stay around for a long time or a short amount of time." 00:12:55 This is a bet that unified primitives beat tiered ones.
A hyperscaler deliberately leaves GPU revenue on the table
Conventional logic says sell every scarce GPU to the highest bidder. Garman says AWS could place all its accelerators with frontier labs and chooses not to, treating startups as an investment in the ecosystem and its own future enterprise base.
Garman: "we could sell every single GPU or AI accelerator we had to probably just the big frontier labs and cloud a day. We choose not to do that because we actually want to keep growing the full ecosystem." 00:18:17
The "bubble" is a portfolio problem; concentration is the real risk
Rather than debating whether AI demand is real, Garman reframes it as VC math applied to infrastructure: some customers will fail, but breadth means no single failure matters, and the internet bust still produced durable giants.
Garman: "is every billion dollar startup going to make it? No, they won't but you know that's kind of the game... one makes it and pays for the other 10." 00:22:17 And the dotcom precedent: "a lot of the companies that had durable businesses the Google Amazons the others like that they did pretty well." 00:22:46
Naming and positioning: Trainium is "maybe the best inference chip on the market"
The conventional read is that a chip named for training is a training chip. Garman says AWS's own naming confused the market and that Trainium is arguably the best inference silicon today, with the big Bedrock inference volume already on it.
Garman: "Originally we had a chip called Inferentia for Inference... it turns out that Trainium is actually maybe the best inference chip on the market right now from absolute performance and cost performance point of view." 00:37:47
Powerful models as attack surface and as the solution
Instead of treating frontier model capability purely as a security threat, AWS argues the same models should power defenders at machine speed.
Garman: "that is a real risk I think to customer environments but it's also a real opportunity and so we recently launched a service called Continuum that uses these powerful models to help customers go secure their environment." 00:49:55 "at some point customers are going to need security at machine speed not at human speed." 00:50:25
3. Companies Identified
Amazon Web Services (AWS)
The world's largest cloud provider, about $169-170B in revenue growing 37%. The subject of the episode; the guest's own company. Why mentioned: Still early in its opportunity. Garman: "we're still at the early stages of what the business can be... there's a huge amount of workloads still that live on-prem today." 00:02:20 Revenue: "about 169, 170 billion... 37 percent." 00:02:11
Amazon (parent)
Parent of AWS, funding the $220B capex. Why mentioned: Garman: "the company has been really good about funding that and obviously the AWS business is a good one that we like to invest in." 00:20:18
Anthropic
Frontier AI lab, major AWS customer building on Trainium. Why mentioned: Cited as a core frontier partner and Bedrock growth driver. Garman: "we're great partners with the large frontier labs, the Anthropics and OpenAIs and Meta." 00:17:18 "We have great deals with both Anthropic and OpenAI to build on top of Trainium." 00:36:53 "it's why you see Anthropic really growing really rapidly." 00:45:40
OpenAI
Frontier AI lab and AWS customer. Why mentioned: Garman: "it's why you see open AI workloads migrating to Bedrock." 00:45:40 Also a Trainium partner. 00:37:22
Meta
Large AI customer of AWS. Why mentioned: Listed among the large customers AWS invests in. 00:17:48
NVIDIA
GPU supplier. Why mentioned: Garman: "we recently announced we're going to be buying 2 million NVIDIA GPUs over the next couple of years." 00:00:21
Salesforce
Large enterprise AWS customer. Why mentioned: An example of large enterprises with demand for fewer accelerators, alongside JPMorgan Chase. 00:17:48
JPMorgan Chase (JPMC)
Large financial enterprise customer. Why mentioned: Same large-enterprise customer example. 00:17:48
Firecracker (AWS micro-VM technology)
AWS-invented micro-VM technology adopted by sandbox startups. Why mentioned: Garman: "a lot of the sandbox companies, startups use Firecracker." / "They all use Firecracker, which we invented 10 years ago maybe." 00:15:17 "they have a great security boundary and you don't have a lot of that virtualization overhead." 00:15:17
AgentCore and Bedrock (AWS services)
AWS agent-building platform and managed model service. Why mentioned: Core to AWS's agent and enterprise AI strategy. Bedrock is where "vast majority" of proofs of concept land in production. 00:46:11 "agent core we build these building blocks so it's easier to build agents with any of the models." 00:46:42
Amazon SageMaker
AWS model building and hosting platform. Why mentioned: Resurging for open-weights post-training. Matt Garman (host framing): "SageMaker is getting a new lease of life." 00:48:00
Graviton (AWS Arm CPU)
AWS custom CPU. Why mentioned: "20% cheaper at a 20% better performance." 00:35:52
Trainium (AWS AI accelerator)
AWS custom AI chip, now at generation 3 with Trainium 4 announced. Why mentioned: Sold out into late next year; powers most Bedrock inference. 00:36:53
Amazon Quick
AWS/Amazon agent and enterprise productivity product rolled out to every Amazon employee. Why mentioned: Garman: "we rolled out Amazon quick to every single Amazon employee... that has grown like wildfire." 00:51:12
Continuum (AWS security service)
AI-powered security service that finds and prioritizes vulnerabilities. Why mentioned: Garman: "Continuum is incredibly popular with customers." 00:50:25
Kiro, Claude Code, and Codex (coding agents)
Coding agents customers use to deploy on AWS. Why mentioned: Garman: "they'll tell their coding agent, whether it's Kiro, whether it's Claude, whether it's Codex... I want to build on AWS." 00:09:14
Gemini (Google)
Model usable with AWS agent tooling. Why mentioned: Garman: "you can use Gemini or other things for it." 00:46:42
Neoclouds (as a category)
GPU-focused clouds contrasted with AWS. Why mentioned: Startups choose AWS over a neocloud for security and capabilities, and neoclouds show customer concentration of "30, 40, 50, 60%." 00:06:19, 00:21:17
Google and Amazon (dotcom survivors)
Cited as durable internet-bubble survivors. Why mentioned: Garman: "the Google Amazons the others like that they did pretty well." 00:22:46
Netflix
Used as an illustration of why communities benefit from cloud services. Why mentioned: Garman: "Do you not want to use Netflix and they'll be like no no, I still want Netflix." 00:31:03
4. People Identified
Matt Garman
CEO of AWS; first GM for EC2; interned at AWS in 2005 during business school. Why mentioned: The guest. His intern project analyzed who AWS would be most valuable to: "my project was actually to come up with an analysis of who we thought AWS would be most interesting to. And the answer was startups." 00:03:28
Raghu Raghuram
a16z general partner and former VMware CEO, host of the conversation. Why mentioned: Interviewer who drove the discussion on startups, GPUs, and agents.
Andy Jassy
Amazon CEO and former AWS CEO. Why mentioned: Garman: "Andy's been public about saying this the potential for AWS is really really large and over the next decade the potential is there." 00:23:24
5. Operating Insights
Plan the supply chain four or five tiers deep and guarantee the weakest link
AWS started tracing dependencies a decade ago to find any component that could bottleneck them, not just servers.
Garman: "we saw this problem coming probably a decade ago and really started not just thinking about how many servers do we need to track but just thinking all the way through the supply chain four tiers five tiers down what is the component that could cause an issue for us and making sure we had guaranteed supply on that." 00:28:48
Sell planning as a product: do what customers structurally cannot
Multi-year capacity planning (memory, power, transmission) is a service customers can't replicate and is a core value of the platform.
Garman: "customers can't do themselves they're not going to plan their memory footprint in 2028 they can't do that and so that's one of the values that we spend a huge amount of time thinking about." 00:26:02
Restructure teams into small, mobile pods as agent leverage rises
Instead of fixed 10-person teams per capability, use 3-4 person pods that build fast and then rotate onto new problems while maintaining shipped products.
Garman: "one doesn't have to be 10 people it can be three to four people and they build something so fast you actually want to move them to different projects and problems thinking about how do you both operate and maintain the things that you built while being agile." 00:54:10
Let line-of-business teams build their own agents
Unblocking non-engineers removes the developer queue as a bottleneck.
Garman: "things that used to be blocked by software developers actually the line of business folks are able to go and unblock themselves and innovate more quickly." 00:51:42
Run FDE engagements with a fixed end date and a customer-ownership requirement
Time-box the engagement (about 45 days), teach evals and data labeling, and only engage customers ready to own the result.
Garman: "we want to go into a customer who's ready to accept... really kind of accept owning this when we're done." 00:44:11
6. Overlooked Insights
AWS's quiet admission that coding agents route "default" deploys to partners
In a brief aside, Garman concedes that for brand-new users who aren't on any cloud, agent-driven "just deploy this" requests often flow to AWS partners with simpler layers. This signals a real distribution gap at the very top of the funnel, one the new 30-second Gmail signup is evidently designed to close, and an opening for easy-deploy layers sitting on top of hyperscalers.
Garman: "if you're brand new, you don't already have an AWS, you already haven't set up your IAM... a lot of times that'll go to some of our partners that have an easier to use layer on top, which we love by the way too." 00:09:44
The elasticity of cloud has partially disappeared in the GPU era
Said almost as an aside, this reframes the economics of AI infrastructure: the defining cloud promise of instant, elastic capacity doesn't hold for GPUs, which favors those with allocation, long-term commitments, and planning capability over those who rely on on-demand access.
Garman: "one of the most painful things is that with the real ramp of GPUs like a lot of the elasticity has unfortunately kind of gone away and so hopefully we'll get back to it. And in our core compute and storage and things that elasticity is still there." 00:20:47