BREAKING: “Devin Writes 95% of Our Code” Scott Wu, CEO of Cognition
- 01AI Coding Agents Are Becoming Organizational Infrastructure, Not Just Developer Tools
- 02### The Real Value Is Capacity Expansion, Not Efficiency Gains
- 03### Recursive Self-Improvement Is Already Happening at the Product Level
- 04### Government Is the Next Massive Frontier for AI Coding Agents
- 05### AI Coding Category Is Consolidating Rapidly Through Acquisition, Not Attrition
1. Key Themes
AI Coding Agents Are Becoming Organizational Infrastructure, Not Just Developer Tools
The competitive moat is shifting from model quality to change management and enterprise deployment. Enterprises are making long-term workflow bets, not benchmark bets — meaning the application layer companies that can actually land and expand inside organizations will win regardless of which underlying model leads any given week.
"You don't wanna teach all of your engineers or your entire team how to use one particular suite & one particular product & then find out, oh, actually people aren't using this one anymore because everyone says this other one is better."
"It's not as simple as throw the tool over the wall, here's your chatbot, try that out, hopefully it's good for you. There's a lot of fundamental questions that you have to think through and work through."
### The Real Value Is Capacity Expansion, Not Efficiency Gains
The efficiency framing ("this task now takes a third of the time") undersells the opportunity. The bigger unlock is net-new software that was previously impossible to justify building — a shift analogous to the Industrial Revolution, where mass production arrived before anyone knew what to mass-produce.
"The real unlock is not in efficiency but in capacity. What can we do if we can build so much more?"
"Anything that involves using a computer in some form is ultimately some kind of software, & when all the software drives itself.. what do you want to do? What do you want to create?"
### Recursive Self-Improvement Is Already Happening at the Product Level
Cognition is the most concrete live example of an AI system materially accelerating the development of itself. Devin writing 95% of Cognition's code — and total code shipped growing 7x in six months — is not a future scenario; it is a disclosed operating metric.
"If we only 3x'd year over year, we'd be lagging the market."
"Devin now writes roughly 95% of Cognition's own code, total code shipped has grown 7x in six months, and customer usage has increased 11x to 12x."
### Government Is the Next Massive Frontier for AI Coding Agents
Wu frames the U.S. government as arguably the single biggest untapped market for software development capacity — enormous backlog, chronic engineering shortages, and a growing willingness to buy. Cognition is actively positioning there with customers already at NASA, the U.S. Army, and the U.S. Navy, and held a July 4th event in Washington D.C. as a deliberate signal.
"He describes government as possibly the single largest case of an organization that needs software built and does not have the engineers to build it."
"There will come a day where all the businesses of the world are built on AI, & I think it's something that we should be thinking about sooner rather than later."
### AI Coding Category Is Consolidating Rapidly Through Acquisition, Not Attrition
The market is not waiting for competitive winners to emerge organically. Large acquirers (Google, SpaceX) and well-capitalized independents (Cognition) are buying their way into full-stack positions, compressing the window for standalone tools to remain independent. Windsurf went from a $1.25B independent to absorbed inside 12 months; Cursor was acquired by SpaceX at $60B.
"The category has consolidated through acquisition rather than attrition. OpenAI's $3 billion offer for Windsurf expired in July 2025. Google then paid $2.4 billion to license Windsurf's technology... and Cognition bought what remained days later."
"On 16 June 2026 SpaceX agreed to buy Anysphere, the maker of Cursor, for $60 billion in all-stock consideration."
2. Contrarian Perspectives
Benchmarks Are Meaningless as a Proxy for Real-World AI Progress
The consensus among AI buyers is to track benchmark performance as a proxy for model capability. Wu argues this is structurally misleading: benchmarks are, by definition, pre-specified tasks with defined success criteria — exactly the condition that makes RL trivially solvable. A model "winning" a benchmark tells you almost nothing about its utility on novel, open-ended work.
"The thing that's counterintuitive is we're kind of getting to the point where you can solve basically any benchmark. Because what does it mean to have a benchmark? It means you've already defined the task, you've clarified what success or failure looks like... And the truth kind of is, well, if you do that, then you can teach the model to go do that. RL works."
Evidence: The right metric is actual business outcomes — not tokens consumed, not benchmark scores. Cognition explicitly tracks productivity and ROI rather than model-centric metrics, which Wu calls the metric "every AI company gets wrong."
### The "One Lab Wins Everything" Scenario Is Wrong — Independence Is a Real Strategy
The dominant narrative in AI is that foundation model labs will vertically integrate and capture most of the value. Wu rejects this directly, arguing enterprises will not consolidate on a single lab's toolchain, and that a well-capitalized independent has structural advantages in serving them.
"There's been a narrative out there that in order to win, you have to be a lab yourself, or you have to go sell to a lab. We've always believed there's a lot of value & a lot of power in being independent."
"There's a world that folks sometimes talk about where there's the full recursive superintelligence, & then everything is all owned by a single entity... I just don't think that's the right future for us."
Evidence: Enterprise customers explicitly refuse to standardize on a single lab — creating durable demand for an independent orchestration layer. The internet analogy is instructive: the open internet won over a handful of corporate-controlled networks.
### Token Consumption Is the Wrong Success Metric for AI Products
Conventional wisdom in AI infrastructure is to measure usage and engagement via token consumption. Wu explicitly dismisses this as a vanity metric for coding agents, arguing that outcomes — not activity — are what enterprises actually pay for.
"Obviously literal tokens is just not the right answer. What you care about is outcomes and what you're actually delivering. There's no secret trick or easy way out in terms of measuring actual ROI, productivity, & so on."
Evidence: Cognition's 11–12x growth in customer usage was driven by demonstrable enterprise outcomes (e.g., Mercedes-Benz reducing an 8-month COBOL modernization to 8 days; Itaú auto-remediating 70% of security vulnerabilities).
3. Companies Identified
Cognition Applied AI lab; creator of Devin autonomous AI software engineer and Devin Desktop (formerly Windsurf). Why mentioned: Primary subject. $2.5B raised, $26B valuation, ~$492M ARR, Devin writes 95% of its own code.
"The company has raised more than $2.5 billion, including a $1 billion round in May 2026 at a $26 billion valuation, while growing annualized revenue from roughly $37 million to $492 million in a year."
Windsurf (now Devin Desktop) AI IDE; formerly independent, acquired by Cognition in July 2025 after Google hired its leadership for $2.4B. Why mentioned: Key acquisition case study; $82M ARR, 350+ enterprise customers at time of acquisition.
"Windsurf had built one of the most widely used AI IDEs and a real enterprise business, $82 million of ARR across more than 350 enterprise customers."
Poke / The Interaction Company AI agent reachable via iMessage, SMS, Telegram, WhatsApp. Acquired by Cognition July 23, 2026 for a reported low nine figures. Why mentioned: Second Cognition acquisition; 100M messages in 3 months; first AI agent on Apple Messages for Business.
"Poke launched in March 2026 as an agent reachable over iMessage, SMS, Telegram and WhatsApp, passed 100 million messages in three months, and became the first AI agent approved on Apple's Messages for Business platform."
Anysphere / Cursor AI coding tool; acquired by SpaceX for $60B in June 2026. Why mentioned: Major consolidation data point; market share fell from ~41% to ~26% of corporate AI coding spend in one year.
"On 16 June 2026 SpaceX agreed to buy Anysphere, the maker of Cursor, for $60 billion in all-stock consideration."
Founders Fund Venture capital firm. Why mentioned: Led Cognition from seed through Series D; seed ($21M), Series A ($175M at $2B), Series C ($400M at $10.2B).
"Founders Fund has backed Cognition from seed through Series D, leading most of its financings."
Mercedes-Benz Automotive manufacturer. Why mentioned: Pilot case study — Devin analyzed 200,000+ lines of COBOL and reduced an estimated 8-month modernization to 8 days.
"A 4-week Mercedes-Benz pilot had Devin analyze more than 200,000 lines of COBOL and reduce an estimated 8-month modernization project to 8 days."
Itaú Brazilian bank. Why mentioned: Production deployment case study — Devin auto-remediates 70% of identified security vulnerabilities.
"Devin automatically remediates 70% of identified security vulnerabilities at Itaú."
Goldman Sachs, Citi, Santander, Nubank Financial services firms. Why mentioned: Named production customers of Devin in banking.
"Customers include Goldman Sachs, Citi, Santander, Nubank & Itaú in banking."
NASA, U.S. Army, U.S. Navy U.S. government agencies. Why mentioned: Active government customers; signal of Cognition's public sector push.
"Government agencies including NASA, U.S. Army & U.S. Navy."
SpaceX Aerospace company. Why mentioned: Acquirer of Cursor at $60B — the most significant consolidation event in the AI coding category to date.
"SpaceX agreed to buy Anysphere, the maker of Cursor, for $60 billion in all-stock consideration."
Lux Capital, General Catalyst, 8VC Venture capital firms. Why mentioned: Co-led Cognition's $1B+ Series D at $25B pre-money.
"Its latest financing, a $1 billion+ Series D announced on May 27, 2026, was co-led by Lux Capital, General Catalyst, and 8VC."
4. People Identified
Scott Wu Co-Founder & CEO, Cognition. International Olympiad in Informatics gold medalist. Why mentioned: Primary interview subject; architect of Cognition's strategy, the Windsurf acquisition, and the independence thesis.
"There's been a narrative out there that in order to win, you have to be a lab yourself, or you have to go sell to a lab. We've always believed there's a lot of value & a lot of power in being independent."
Walden Yan Co-Founder, Cognition. IOI gold medalist. Why mentioned: Had the "eureka moment" for Devin — handed an unsolvable MongoDB error to his personal agent prototype, which fixed it autonomously in December 2023.
"I couldn't sleep that night. We were just like, shit, maybe it does work. Maybe you can just have an AI engineer buddy that can just do work for you."
Steven Hao Co-Founder, Cognition. IOI gold medalist. Why mentioned: Co-founder; built one of the original internal agent prototypes ("dev Steven") that became Devin.
"Each founder built an agent version of himself, a dev 'Steven' & a dev 'Walden,' and the ideas were eventually folded into one product."
Napoleon Ta Partner, Founders Fund (leads growth practice); sits on Cognition's board. Why mentioned: Lead investor from seed through Series D; offered to play heads-up poker to negotiate the seed terms.
"Since you're such a big fan of poker, we could just play a heads-up poker match for that, and we could have the winner decide whose terms we actually sign. In retrospect, it would've been for 1% of the company."
Peter Thiel Co-Founder, Founders Fund. Why mentioned: Vetoed the poker negotiation; noted as a formidable chess player Wu explicitly refuses to challenge.
"I would not recommend playing against Peter Thiel in chess. You gotta pick games that you can win."
Jeff Wang Interim CEO / Head of Business, Windsurf (post-Google acquisition); now leads Windsurf product line as a division of Cognition. Why mentioned: Key principal in the 48-hour Windsurf acquisition; described the post-Google all-hands mood as "bleak."
"Wang held an all-hands that Friday where most of the team expected to hear that OpenAI had bought them, and instead heard about Google."
Graham Moreno Interim President / VP of Global Sales, Windsurf; now leads alongside Wang at Cognition. Why mentioned: One of the four principals who worked out the Windsurf deal over a single weekend.
"Wu & Russell Kaplan reached Wang that evening, and the deal was worked out between the four principals across the weekend."
Russell Kaplan Cognition team member. Why mentioned: Named as one of the four principals in the Windsurf weekend negotiation.
"Wu & Russell Kaplan reached Wang that evening, and the deal was worked out between the four principals across the weekend."
5. Operating Insights
Acquire for GTM and Distribution, Then Let Product Integration Follow Usage Organically
Cognition's Windsurf integration playbook rejects the conventional 30/60-day forced integration timeline. They ran joint offsites, rebuilt billing, and found shared office space — but let actual user behavior and new feature development determine when and how the products merged technically. The result: a culturally unified team within a year, with Windsurf's GTM org (350 enterprise customers, $82M ARR) intact.
"Rather than force it & say, over the next 30 days or 60 days we have to integrate these products, it was much more letting the actual usage & the new features that we were building pull the products together more gradually rather than forcing it."
"A lot of our team are former founders. Of our first 56 people, 30 of us had founded a company before this. A lot of us together were people who liked being ambitious, entrepreneurial, & thinking from first principles."
### Measure AI ROI by Business Outcomes, Not Activity Metrics
Teams deploying AI coding agents should resist the temptation to track tokens, completions, or acceptance rates as KPIs. The metric that matters — and the one enterprise buyers will ultimately pay for — is actual productivity and business output. This reframes the sales conversation and the internal evaluation framework.
"Obviously literal tokens is just not the right answer. What you care about is outcomes and what you're actually delivering. There's no secret trick or easy way out in terms of measuring actual ROI, productivity, & so on."
### Rolling Out Coding Agents Requires Rethinking the Entire Engineering Workflow, Not Just Tooling
Operators adopting AI coding agents should expect to redesign how teams write specs, conduct user research, plan sprints, and manage engineering handoffs — not just swap out the IDE. Companies that treat it as a tooling change will underutilize it; companies that treat it as a workflow redesign will 10x output.
"For Cognition, the larger challenge is organizational rather than technical. Rolling out AI coding agents changes how teams plan work, write specifications, conduct user research and manage engineering workflows. Success depends as much on implementation as on model capability."
6. Overlooked Insights
Cursor's Market Share Has Already Peaked and Is in Decline
Despite Cursor's $60B SpaceX acquisition price, its actual share of corporate AI coding spend — tracked via Ramp card data — fell from ~41% in June 2025 to ~26% by May 2026. The acquisition may be buying a declining share position, and suggests the AI coding tool market is more fragmented and contested than the headline valuation implies.
"Cursor had reached roughly $4 billion of annualized revenue before the SpaceX agreement. Its share of corporate AI coding spend, measured by Ramp card data, fell from around 41% in June 2025 to about 26% by May 2026."
Devin's Origin Was a Single Unsolved Database Error — The Minimum Viable Demonstration for Agentic AI Is Surprisingly Low
The moment that made the Cognition founders believe autonomous AI engineering was possible was not a complex system — it was an agent fixing a MongoDB installation error that a human engineer had given up on. For entrepreneurs evaluating whether agentic AI is ready for their use case, the signal is that the bar for a compelling proof-of-concept is a task that is annoying, narrow, and well-defined — not impressive or broad.
"You need encyclopedic knowledge of all of the different errors that you could run into and what would cause each of them. And then you need the ability to actually run and diagnose things. Those 2 things, the encyclopedic knowledge & the ability to actually go & run commands, that's literally what a coding agent is."