How Kavak Rebuilt Itself Around AI Agents | Alejandro Maza Ayala
- 01From Transactional to Relational: The Agent-Per-Customer Architecture
- 02Evals as the Engine of Speed, Not a Safety Tax
- 03Superhuman Agents as the Actual Performance Target
- 04Destroying Working Architecture to Capture a New Paradigm
- 05The AI CEO Experiment: Carved Out a City as a Testbed
- 06The Jedi Academy: Company-Wide AI Retraining from CEO to Mechanic
1. Key Themes
From Transactional to Relational: The Agent-Per-Customer Architecture
Kavak made a foundational shift from measuring transactions to managing long-term customer relationships at scale. Rather than routing customers through task-specific workflows, they instantiate a dedicated agent per customer that holds memory across years of interactions and pursues a single long-term goal.
"When a customer comes in right now, an agent will get spawned specifically for this customer with its own virtual machine. It will remember years of interaction of these customers with Kavak, what they visited in the webpage or a call they had two years ago. Remember everything in its memory. Come up with a strategy and set a long-term goal to maximize the lifetime value of this customer." 00:04:00
Evals as the Engine of Speed, Not a Safety Tax
Rather than treating evaluation as a compliance step, Kavak treats it as the primary enabler of velocity. The rule of thumb they apply is striking in its resource commitment.
"A good rule of thumb here is we spend about the same amount of time, engineer time, tokens, and money on building the evals than building the agents. And this is how you get better and better and better. Not letting evals as an afterthought." 00:09:10
Superhuman Agents as the Actual Performance Target
Kavak explicitly set the benchmark not at "good enough" but at outperforming the best human ever hired on every dimension that matters. The results confirmed the bet.
"We tripled NPS and customer satisfaction score by putting the agent in front of the customer. And at first it converted like 50% more than our human team. And now it's converting over that like 2.1X more." 00:12:19
Destroying Working Architecture to Capture a New Paradigm
One of the most consequential decisions was tearing down a multi-agent system that was already generating profitability and growth, because a new model release revealed the prior architecture would become a ceiling rather than a foundation.
"Claude Opus 4.5 came out and I realized this isn't the right paradigm anymore. Like the intelligence now doesn't need the graph and the multi-agent latticework and harness because it will constrain this level of intelligence. So we decided to destroy everything we had been building for two years that was working, that brought us to profitability, that brought us amazing growth." 00:28:53
The AI CEO Experiment: Carved Out a City as a Testbed
Rather than theorizing about AI leadership, Kavak ran a live experiment — assigning an agent as CEO of an entire city operation.
"We carved out a city in Mexico. It's Cuernavaca. And we put an agent in one of our harnesses as a CEO. And it starts learning and it starts making decisions and evaluating on those decisions. It's only been running for six weeks now. The goal of the first month was to double the profits of Cuernavaca. It didn't reach it, but it was 1.5x, like 50% more profits just by managing the city." 00:16:30
The Jedi Academy: Company-Wide AI Retraining from CEO to Mechanic
Rather than hiring new talent into AI roles, Kavak redesigned their own workforce by building an internal academy where every employee — regardless of role — learns to build and collaborate with agents.
"We launched a program inside Kavak that's called the Jedi Academy where anyone from Kavak, from the CEO to AI engineers to mechanics, goes to the academy. It's super hard. And after six weeks, they launch state-of-the-art AI agents to production. And it's mechanics and finance guys and engineers, like everyone can do it." 00:00:39
The Token Tier Framework: Measuring ROI of Every Token Spent
Alejandro introduced a framework for thinking about AI spend quality that cuts through vanity metrics around adoption.
"Tier three tokens, the most valuable, are these agents where you can get the ROI of each specific token. Tier two tokens are things that you can measure indirectly. Do I see improvements in the code base? Tier one, when most companies are, is people are just using GitHub Copilot or ChatGPT or whatever. What happened with those? I have no idea." 00:26:59
The Electricity Analogy: Creative Destruction at Industrial Scale
Alejandro frames the current AI moment using a precise historical parallel — Edison's dynamo and Ford's factory — to explain why superficial adoption yields only marginal improvement while full architectural redesign yields 3x-plus gains.
"What needed to be done was to destroy that factory, build it on a flat surface, and redesign your whole factory around small dynamos and electricity. And then you get like the 3x improvement in productivity that powered the U.S. during the 20th century. The same is happening again today. People want to adopt it, but they're not willing to redesign the whole company. And they just adopt it superficially. And in the end, that'll give you a 6% or a 10% improvement, not a 10x improvement." 00:33:32
Human-in-the-Loop Must Close the Feedback Loop, Not Just Escalate
A key architectural insight is that the standard escalation model — agent fails, kicks to human support, forgets — is broken. The correct design keeps the agent obsessed and learning.
"Usually, if an agent hits a wall or can't perform anymore, it'll send this case or this customer to a tier two support and forget about it. That doesn't really work. Because you don't close the loops, so you don't generate the data to train the agent to do this better." 00:24:09
2. Contrarian Perspectives
Customers Prefer Buying Expensive, High-Stakes Products From AI Over Humans
The conventional assumption is that high-ticket, trust-sensitive purchases (used cars, personal loans) require human touch. Kavak's data contradicts this entirely.
"We never built customer support or customer service agents. We built sales agents. The experience for the customer is amazing. We tripled NPS and customer satisfaction score by putting the agent in front of the customer. And at first it converted like 50% more than our human team. And now it's converting over that like 2.1X more." 00:10:57
The CEO Is Not the Last Job AI Takes — It May Be Among the First
Conventional wisdom holds that creative, strategic, and leadership roles are the final frontier for AI. Kavak's Cuernavaca experiment suggests otherwise.
"It's only been running for six weeks now. The goal of the first month was to double the profits of Cuernavaca. It didn't reach it, but it was 1.5x, like 50% more profits just by managing the city, which is crazy, right? And it's the CEO. Like people were like, that was the last job AI was supposed to take. And no, it isn't really." 00:16:30
Working AI Systems Should Be Destroyed When a Better Paradigm Emerges
The standard operating logic is: if it's working and generating growth, don't touch it. Alejandro argues the opposite — locking into a working architecture is the mistake.
"We decided to destroy everything we had been building for two years that was working, that brought us to profitability, that brought us amazing growth. And start over with a harness that we thought would be robust and scalable and leverage recursive self-improvement or new models, more intelligent models coming out every month." 00:29:14
Large Company AI Adoption Is Structurally Incapable of Deep Transformation
Rather than being optimistic about enterprise AI adoption, Alejandro is blunt: the incentive structures of public companies make genuine transformation nearly impossible, which is the opportunity signal for new founders.
"It's really hard for a CEO today, especially of a large company or public company to go and say, I'm betting everything on AI. Like how many CEOs will do that in a company at scale? So while they adopt, new companies can be formed that are built around the strengths of AI and take over." 00:32:20
AI Transformation Must Be Top-Down Mandated, Not Bottom-Up Adopted
The popular playbook of hackathons and grassroots AI champions is dismissed as ineffective.
"I've seen so many companies. It's just like, oh, we're doing a hackathon. People are coming up with use cases. We're sponsoring some of these use cases. That doesn't work. Be very clear on what the company will look like in three or five years and then start building that... An army doesn't really work if everyone comes up with ideas on the strategy and tactics and does whatever they want." 00:26:07
3. Companies Identified
Kavak Used car marketplace and vertically integrated automotive platform in Latin America (buying, refurbishing, selling, financing, insurance). Mentioned as the primary subject — a company that has rebuilt itself entirely around AI agents, achieving 96% of all interactions handled by agents, 2.1x conversion improvement over human sales teams, and 50% profit improvement in one city under an AI CEO.
"96% of all interactions are handled by agents. 95% of all transactions are completely handled by agents. Every day, between 100 and 200,000 agents get instantiated in a day." 00:08:14
Opi Analytics Pre-ChatGPT machine learning company founded by Alejandro Ayala. Served 1,500 companies across risk algorithms, logistics, forecasting, and marketing. Mentioned as context for Alejandro's deep pre-LLM AI background.
"We served 1,500 companies around like risk algorithms, logistics, forecasting, marketing. But really the power of Transformers and then the ChatGPT moment when it arrived made things very clearly that we could now build a whole new company and way of building companies." 00:02:45
Anthropic (Claude) AI model provider. Claude Opus 4.5 was the specific model release that triggered Kavak's complete architectural rebuild — the catalyst for moving from multi-agent graphs to the per-customer virtual machine harness.
"Claude Opus 4.5 came out and I realized this isn't the right paradigm anymore." 00:28:53
4. People Identified
Alejandro Ayala (Ale) Chief Product and AI Officer at Kavak; founder of Opi Analytics. Architect of Kavak's full AI-native transformation — designed the Jedi Academy, the per-customer agent harness, the AI CEO experiment in Cuernavaca, and the token tier ROI framework. Mentioned throughout as the primary builder of one of the most advanced real-world agentic deployments at scale.
"We bet the company on transforming to a company run by agents. The questions we ask ourselves is how would we build Kavak in 2035 with GPT-10 level intelligence?" 00:04:00
Carlos (Kavak CEO) Co-founder/CEO of Kavak. Mentioned briefly as the leader who brought Alejandro in to build the AI transformation.
"We joined Kavak and Carlos to build it." 00:02:45
Joseph Schumpeter Economist. Cited by Alejandro as the intellectual foundation for why AI will create value through destruction of incumbents rather than adoption.
"There's this concept in economics about creative destruction from Joseph Schumpeter. And what it says is that the way innovation hits the economy isn't by companies adopting the new technology, but by companies remaining the way they were and incumbents with a new technology destroying the old companies." 00:31:25
Thomas Edison Inventor. Cited in the electricity/dynamo analogy to explain why superficial technology adoption (swapping one engine for another) yields only marginal gains versus full architectural redesign.
"Edison started commercializing electricity in New York and then London, and he invented a dynamo that was extremely efficient. So you could have built Ford's factory 40 years before Ford. The technology was there." 00:33:05
Henry Ford Industrialist. Used as the historical example of an entrepreneur who built a new architecture from scratch around electricity rather than retrofitting the old one — the model for what AI-native founders should do today.
"The technologies for Ford's production line were developed in 1879 and 1881... What needed to be done was to destroy that factory, build it on a flat surface, and redesign your whole factory around small dynamos and electricity." 00:33:05
5. Operating Insights
Evals Must Be Business-Outcome-Anchored, Not Activity-Anchored
Most companies measure AI performance on proxy metrics (call length, number of interactions, response speed). Kavak's eval framework cuts straight to commercial outcomes.
"First and foremost, the results for the business. If my customer is happy, they'll buy a car. They'll get their loan approved. They'll sell a car to us. And that's the first check. Did it convert? I see companies measuring number of calls or minutes during the call or some superficial KPIs that give you some information, but that doesn't really work." 00:09:38
Never Let a Failed Agent Interaction Escape the Feedback Loop
The specific operational failure mode to avoid: when an agent can't complete a task and escalates to a human, that escalation must feed back to the agent — not disappear into a support queue. Every unresolved case is a training data point being abandoned.
"If an agent hits a wall or can't perform anymore, it'll call this API saying, I need help. And on the other side, it's not an agent or software. It's a human helping them out... You don't close the loops, so you don't generate the data to train the agent to do this better." 00:24:09
Rebuild APIs First Before Deploying Agents
Before any agent can perform, the underlying system infrastructure must be designed for agents to use — this is the prerequisite step most companies skip, which is why they see no efficiency gains.
"You need to redesign your whole company around the agents and around the future capabilities. And this means really rebuilding most of your APIs, rebuilding your system so the agents can use them to perform." 00:05:41
Move Fast by Investing in Brakes, Not by Going Slow
The counterintuitive operating principle: the way to deploy AI at speed without breaking things is to invest heavily in evals upfront — not to slow down rollout.
"I like to move extremely fast. But in order to move fast, you need to have brakes, right? Imagine a car. You'll hit on the gas just if you have the right brakes... How fast can we go? Well, it depends on the quality of our evals." 00:08:43
6. Overlooked Insights
Vertical Integration Becomes an AI Superpower in Emerging Markets
This was mentioned almost in passing as background context, but it is deeply significant for investors: Kavak had to build every piece of infrastructure (fintech, logistics, vehicle history data) because none existed in Latin America. This vertical stack — which looked like a burden — is now the exact reason their agents can underwrite loans in three minutes, personalize financing, and manage full customer journeys. The data moat created by forced vertical integration is what makes the agent architecture possible at all.
"To do that, we also had to build a fintech and a logistics company and the Carfax and basically all the infrastructure for this to work didn't exist in LATAM. So we had to build everything vertically so we could serve our customers the right way." 00:03:22
The implication: any company that has been forced into deep vertical integration in emerging markets — even for painful operational reasons — may now hold an AI data advantage that pure-play horizontal companies cannot replicate quickly.
Warranty Reduction as a Hidden Quality Signal for Agent-Assisted Physical Work
The "El Mike" sidekick for mechanics was mentioned briefly, but the outcome was buried: warranties fell 20-26% after deployment. This is not a customer service metric — it is a product quality and unit economics metric. Reduced warranty claims directly impact margins on physical goods at scale, and this was achieved not by replacing mechanics but by giving them an AI collaborator.
"Warranties came down around like 20, 26% since we launched. And customer satisfaction, again, went up." 00:18:46
The non-obvious insight: AI applied to physical inspection and repair work — not just white-collar knowledge work — can materially improve product quality and reduce after-sale liability. This is an underexplored application area for AI in hardware, automotive, manufacturing, and logistics businesses.