Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/THE AI CORNER/Can AI Agents Actually Use a Com…
NEWS
// NEWSLETTER ISSUE
THE AI CORNER

Can AI Agents Actually Use a Computer? Here's the Real Answer

DATE October 8, 2026SOURCE THE AI CORNERPARTICIPANTS THE AI CORNER
// KEY TAKEAWAYS6 ITEMS
  1. 01Theme: The verification bottleneck has replaced capability as the real constraint
  2. 02Models now beat the human baseline on computer-use benchmarks
  3. 03The hard part is now knowing whether the work was done correctly
  4. 04Theme: Benchmark scores and business outcomes are different units
  5. 05Per-step reliability compounds into poor end-to-end reliability
  6. 06Headline scores aren't comparable across vendors
// SUMMARY

1. Key Themes

Theme: The verification bottleneck has replaced capability as the real constraint

Models now beat the human baseline on computer-use benchmarks

The capability question is settled. The article frames the crossover as real and recent.

"Claude Fable 5 ranked top of OSWorld-Verified with 85%, sourced to llm-stats in June 2026. Human testers score around 72% on the same task set."

"Last year's best model sat at 58% of human performance and this year's is at 119%. Something real crossed over in early 2026."

The hard part is now knowing whether the work was done correctly

The article's central claim is that finishing a task and completing it correctly are different, and agents can't tell the difference.

"The agent did everything right. It read the screen, filled every field and clicked the right buttons, and it still had no way to learn that correct and finished are different things."

"The one that replaced it is whether you can tell when they got it wrong, cheaply enough and fast enough for it to matter. Where you can, the economics already beat every human alternative and the deployments are real and growing. Where you cannot, a smarter model changes absolutely nothing, because the model was never the thing that failed."

Theme: Benchmark scores and business outcomes are different units

Per-step reliability compounds into poor end-to-end reliability

Real office jobs are chains of steps, so small per-step failure rates multiply into large job-level failure rates.

"So if a job has 20 steps and each step works 95% of the time, the job works 36% of the time. That is 0.95 multiplied by itself twenty times."

"Going from 95 to 98 per step nearly doubles end-to-end success at 20 steps, and still leaves you failing two thirds of the time at 50."

"Model improvements arrive in a straight line, but failure compounds. That is an unfair fight, and it is why the last few points on a leaderboard matter far less at work than they do in a launch."

Headline scores aren't comparable across vendors

Benchmark numbers bundle the model with scaffolding, step limits, and test conditions.

"So an 85 belongs to a model, plus the software wrapped around it, plus how many steps it was allowed to take. Two companies both claiming 85 may not be running the same thing at all."

"Some results get re-run by the people who maintain the leaderboard and others are simply reported by the labs, using different time limits, different operating systems and different tools."

The model is becoming a swappable commodity underneath the product

Heavy users no longer care which model does the work.

"One of them turned out to be running millions of automated tasks a month without knowing which model does the work. He had never needed to find out. His supplier swaps models underneath him the way a cloud provider swaps servers."

"When the heaviest users stop checking which model they are on, the model has stopped being the thing that decides whether the work gets done."

Theme: The winning architecture is old-school deterministic code, with AI as the repair crew

Run-caching converts many stochastic steps into one checkable event

The agent runs a workflow once, the system caches it as code, and the model is called back only when something breaks.

"The agent runs a workflow once, the system caches it as deterministic code, every subsequent run executes as cheap code, and the model is only called back when something breaks so it can diagnose, fix and re-cache."

"But it's more like reliability architecture, and what it does is collapse N stochastic steps into one checkable event. You basically stop rolling the dice 50 times and start rolling it once, at the moment the portal changes."

The flagship success case is really an RPA-repair business

The best-cited deployment uses hand-coded scrapers, with the agent maintaining them.

"Their flagship case, reported first-hand, is a CPG data platform running 15 to 20 million automated portal interactions a month, which halved the engineering team assigned to scraper maintenance."

"That is not an agent using a computer. That is an agent fixing the thing that uses the computer."

"If the honest story right now is that AI models are the fix for that 15-year-old flaw, that is a genuinely valuable business worth building. It is also a much smaller claim than the one printed on the tin."

Theme: The cost and error metrics that matter don't exist yet

Cost per verified completed task is the only number that matters, and no one publishes it

Hourly agent pricing ($6-8) ignores retries, escalations, and human cleanup.

"A completed, correct task is" what a buyer purchases. "...cost per verified completed task, including escalation, is the only figure that decides whether any of this works, and I cannot find a single vendor publishing it."

"You land near $2.10 per completed task. Double the headline. And that's the optimistic version, because it assumes every failure gets caught."

Silent errors are the unmeasured killer

Loud failures cost a retry. Silent ones cost relationships and audit exposure.

"An agent transcribing payment terms into an ERP reads net 60 as net 30. The record looks plausible as it passes every visual check. It surfaces weeks later when an invoice goes out wrong."

"A loud failure costs you a retry. A silent one costs you the invoice, the relationship and the audit. Those are not the same 15%, and no benchmark on earth distinguishes them."

Theme: Security is a structural problem, not a bug to be patched

The product requires the exact conditions that make prompt injection dangerous

Agents holding credentials, reading untrusted pages, and acting externally combine all three risk factors by design.

"The recurring finding across serious prompt injection incidents is one configuration where an agent has access to private data, exposure to untrusted content, and the ability to act externally."

"All three conditions. By design. As the product."

Institutions are formally flagging the risk

"The OWASP Top 10 for Agentic Applications, released in December 2025, ranks Agent Goal Hijacking as the number one risk."

"In May 2026 the Five Eyes intelligence agencies published joint guidance naming prompt injection as a main route for manipulating these systems, and told organisations to assume their agents will sometimes behave in ways nobody planned for."

Silent errors and successful injections share the same detection blind spot

"The conditions under which a silent error goes undetected are precisely the conditions under which a successful injection goes undetected. Same gap in the process, wildly different bill at the end."

2. Contrarian Perspectives

Contrarian view 1: The labor data contradicts the "agents are replacing back-office workers" narrative

If $7/hour agents truly substituted for $10/hour offshore labor, offshore employment would be falling. It isn't.

"If agents at $7 an hour genuinely substituted for offshore labour at $10, the offshore numbers would be falling. They are not, and that disagreement is the most useful data in this entire debate."

Evidence: "Global contact-centre employment is forecast to grow from 15.3 million in 2025 to 16.8 million by 2029, even while automation suppresses an estimated 1.9 million individual roles." The Philippines "produced $40 billion in BPO export revenue with 1.9 million workers in 2025 and is tracking toward $42.3 billion and 1.96 million in 2026." Sam Altman said in May 2026 "that he was delighted to have been wrong about customer support roles disappearing," while Vinod Khosla predicted "IT and BPO services will disappear, almost certainly within the next five years." The author declines to pick a winner: "I do not know which, and I would be suspicious of anybody who claims they do."

Contrarian view 2: "Context is the moat" is weaker than it sounds; being the approved vendor is the real moat

The popular claim that accumulated process knowledge creates defensibility doesn't hold against startup competitors.

"Context is the moat. That means the written procedures, the unwritten know-how, who to escalate to. Against a model provider that defends beautifully. Against another startup it defends less well."

Evidence: "If the knowledge lives in one recorded video of somebody doing the job once, that video is cheap to make and any competitor can use it. Procedures and logins belong to the customer, not to whoever holds them this quarter." The author calls this "a switching cost. It makes leaving annoying rather than impossible, and it gets weaker as setting up a new supplier gets cheaper." The more durable moat: "Being the vendor an enterprise is permitted to run in production. Security review, audit trail, liability terms, procurement sign-off. Slow to earn, slow to lose, and impossible to show off on a stage."

Contrarian view 3: The "AI-native" architecture isn't new, and the reframed story is a smaller business than advertised

The architecture serious teams converge on is a rebrand of RPA.

"The design pattern serious teams advocate for is being seen as an emerging AI-native idea. But that is not the case."

Evidence: "software that clicks through screens on a fixed script has existed for 15 years, under the name Robotic Process Automation, and it always broke the moment a website changed." Speed fixes also carry hidden costs: "Reading a page's underlying structure instead of looking at a picture of it is faster, and it ties you straight back to how that particular application is built... Speed bought with fragility is not free."

3. Companies Identified

Anthropic (Claude Fable 5 / Claude Mythos 5)

  • Description: Frontier AI lab with leading computer-use models.
  • Why mentioned: Top performer on OSWorld-Verified; evidence of the capability crossover.
  • Quotes: "Claude Fable 5 ranked top of OSWorld-Verified with 85%"; "with Claude Mythos 5 and Claude Fable 5 tied behind at 85."

Qwen (Qwen3.8 Max)

  • Description: Model from a non-US lab.
  • Why mentioned: Overtook the Claude models on the benchmark within eleven weeks.
  • Quotes: "BenchLM's tracking of the same benchmark had Qwen3.8 Max in front at 86.1%." / "Eleven weeks. New leader, different lab, different continent."

Andreessen Horowitz (a16z)

  • Description: Venture firm that published an August piece on computer-use agents.
  • Why mentioned: Primary source for the cost comparisons, run-caching architecture, and operator interviews, which the author critiques.
  • Quotes: "Andreessen Horowitz published a piece in August, where they interviewed the people actually running these systems." / "The a16z cost comparison puts an agent at $6-8 an hour."

Plaid

  • Description: Fintech and identity infrastructure company; sponsor of the issue.
  • Why mentioned: Sponsored white paper on fraud and identity verification in the agent era.
  • Quotes: "An agent that can file a legitimate claim can file a fake one just as fast. Most identity checks still assume a human at the keyboard." / "Why financial behavior offers stronger, harder-to-fake context."

CPG data platform (unnamed, from a16z)

  • Description: Consumer-packaged-goods data company running agent-maintained scrapers.
  • Why mentioned: Flagship case study of run-caching at scale.
  • Quotes: "a CPG data platform running 15 to 20 million automated portal interactions a month, which halved the engineering team assigned to scraper maintenance."

Brave / Perplexity Comet

  • Description: Brave's security team; Perplexity's agentic browser.
  • Why mentioned: Demonstrated an injection attack via hidden text against a screenshot-driven agent.
  • Quotes: "Brave's security team demonstrated it against Perplexity Comet using white text on white backgrounds and instructions buried in HTML comments, which got the agent fetching one-time passwords out of email."

OWASP

  • Description: Application security standards organization.
  • Why mentioned: Ranked Agent Goal Hijacking the top agentic risk.
  • Quotes: "The OWASP Top 10 for Agentic Applications, released in December 2025, ranks Agent Goal Hijacking as the number one risk."

OpenAI

  • Description: AI lab (led by Sam Altman).
  • Why mentioned: Source of the prompt-injection explainer graphic; Altman's reversal on support jobs.
  • Quotes: "Image source: OpenAI"; Altman "was delighted to have been wrong about customer support roles disappearing."

4. People Identified

Sam Altman

  • Description: CEO of OpenAI.
  • Why mentioned: Reversed his earlier prediction about customer support job losses, a data point against rapid displacement.
  • Quotes: "Sam Altman said in May 2026 that he was delighted to have been wrong about customer support roles disappearing, reversing his own position from ten months earlier."

Vinod Khosla

  • Description: Prominent venture investor.
  • Why mentioned: Represents the bullish-displacement view on IT and BPO.
  • Quotes: "IT and BPO services will disappear, almost certainly within the next five years."

Ruben Dominguez

  • Description: Author of the newsletter piece.
  • Why mentioned: Original analysis supplementing the a16z material; links to a related piece on layoffs.
  • Quotes: "None of what follows appears in their piece." / Related article: "Every Tech Company Is Falling Into A Layoff Trap."

5. Operating Insights

Build for cost per verified completed task, not cost per hour

Measure and price your own deployments on completed, correct outcomes including retries and human escalation.

"Let's assume $7 per hour, 9 minutes per agentic task, gives you about $1.05 an attempt... You land near $2.10 per completed task. Double the headline."

"I cannot find a single vendor publishing it."

Publishing an honest figure here is both a differentiator and a diligence tool when evaluating vendors.

Cache workflows into deterministic code and use the model only for repair

Cutting the number of stochastic steps per run is the highest-leverage reliability move.

"You basically stop rolling the dice 50 times and start rolling it once, at the moment the portal changes."

The proof point: the platform running 15 to 20 million monthly interactions halved the engineering team assigned to scraper maintenance.

Invest in the verification layer and the enterprise-trust stack

Instrument for silent errors, and treat security review, audit trails, and liability terms as product features.

"The firms being valued upward are the ones building the checking layer. That includes supervision, quality control, and handling the cases that go wrong."

"Being the vendor an enterprise is permitted to run in production. Security review, audit trail, liability terms, procurement sign-off. Slow to earn, slow to lose."

6. Overlooked Insights

Unsettled liability and site-permission issues stall deals and could invite pushback

These legal and permission questions are raised quickly but could shape deal velocity and defensibility.

"The first is who pays when an agent files the wrong thing with a regulator? Is it the supplier, the buyer, or the lab that built the model? Nobody has settled that, and it stalls deals in legal review long after the trial run went well."

"The main use case here is automating someone else's website, and plenty of them forbid automated access in their terms and run software to spot it. Some will push back."

The benchmark itself is fragmenting, which makes year-over-year comparisons unreliable

A newer version of the benchmark exists, and eight tasks can be excluded under the official rules, so even the "settled" capability narrative rests on shifting measurement.

"Eight of the tasks can be left out under the official rules."

"There is also a newer version, OSWorld 2.0, whose scores are not comparable with the older one."