The Harness Is the Product: 6 Decisions That Turn an AI Agent Into a Worker You Can Leave Alone
1. Key Themes
Theme 1: The harness, not the model, is the product
Failures live in the wrapper, not the model
- Agent failures (skipped tests, forgotten rules, constant approval clicks) persist across model upgrades because they originate in the surrounding software.
- Quote: "Swap in a better model and every one of those failures stays put, because they live in the software wrapped around the model: what it gets told, what it keeps, what it's allowed to touch, and who checks the result."
A harness is a six-part system you already own, designed or not
- It covers the loop, tools, context management, persistent state, permissions, and the definition of done. The vendor controls only part of it.
- Quote: "Your vendor owns the inner loop, the built-in tools, and the memory handling, and you can change little of it. Everything else is yours: the instructions file, the tests, the permissions, and the definition of done."
- Quote: "The model is the engine, and the harness decides whether you can leave the room."
Theme 2: Three radically different harness architectures all reached production-grade agents
Platform, repository, and role split
- DoorDash, OpenAI, and Anthropic each chose a different shape, which suggests the principle matters more than the specific architecture.
- Quote: "DoorDash, OpenAI, and Anthropic built three completely different harnesses (a platform, a repository, a role split) and got production-grade agents out of all three. Every piece was built by the team running the agent."
- Quote: "Three completely different shapes, and the same lesson under all of them: the team using the agent built the part that made it work."
Theme 3: Agent economics: harnesses pay off on work you couldn't otherwise delegate
Cost multiples buy reliability, not speed
- Anthropic's data shows a harness can cost over 20x more but produce the only working result.
- Quote: "A solo agent built a retro game maker in about 20 minutes for $9, and the central feature was broken. The full three-agent harness took 6 hours and $200, and the game worked."
- Quote: "So the harnessed version cost over twenty times more and produced the only version that worked, which is the trade every vendor demo leaves out."
- Quote: "A harness pays off on work you couldn't hand off at all, and loses money on work where you wanted to save twenty minutes."
Theme 4: Enforce with code and structure, not prose
Deterministic scaffolding beats long instruction files
- OpenAI's giant AGENTS.md failed; the fix was a short index plus linters. DoorDash blends agent steps with deterministic code and a permissioned gateway.
- Quote: "Their first attempt, one giant AGENTS.md file, failed in the way every giant instruction file fails: stale rules piling up until nobody, human or model, read them."
- Quote: "The fix was a roughly 100-line AGENTS.md that works as a table of contents into a structured docs folder, plus architectural rules enforced by custom linters instead of prose."
- Quote: "The work itself is written as YAML playbooks, which DoorDash describes as the Docker container for agent skills, mixing agent steps with ordinary deterministic code."
2. Contrarian Perspectives
1. Upgrading the model won't fix your agent problems
- Against the consensus that better models solve reliability, the author argues the failure modes are structural.
- Evidence: Forgotten rules, unrun tests, and re-prompting persist regardless of model quality.
- Quote: "Swap in a better model and every one of those failures stays put, because they live in the software wrapped around the model."
2. "Getting an agent to write code is mostly solved"
- The hard problem has shifted from code generation to the environment around it.
- Evidence: DoorDash's engineering blog position, backed by 130,000 automated tasks a month and 25,000+ weekly automated code reviews.
- Quote: "The line from their engineering blog worth taping to a wall: getting an agent to write code is mostly solved, and the hard part is the environment around it."
3. Harnesses can be a bad investment
- Counter to the "automate everything" narrative, the author says harnesses lose money on small time-saving tasks.
- Evidence: Anthropic's $9/20-minute vs. $200/6-hour comparison. The teaser also promises a "rule for when a harness costs more than it earns."
- Quote: "A harness pays off on work you couldn't hand off at all, and loses money on work where you wanted to save twenty minutes."
3. Companies Identified
DoorDash
- Description: Food delivery and logistics company with an internal agent platform called Flux.
- Why mentioned: Case study of the "platform" harness.
- Quote: "In a single month in 2026, Flux automated 130,000 engineering tasks, and it now runs more than 25,000 automated code reviews a week across 300-plus reusable playbooks."
- Quote: "Every call to an internal system passes through one MCP gateway that grants only the permissions a job declared and logs everything."
OpenAI
- Description: AI lab behind Codex and the harness engineering post.
- Why mentioned: Case study of the "repository" harness and co-originator of the term's popularity.
- Quote: "Three engineers, later seven, merged roughly 1,500 pull requests, about 3.5 per engineer per day, and throughput went up as the team grew."
- Quote: "The term took off in February 2026, when Mitchell Hashimoto used it on his blog and OpenAI published its harness engineering post a few days later."
Anthropic
- Description: AI lab; published a planner/builder/evaluator multi-agent harness.
- Why mentioned: Case study of the "role split" harness, notable for publishing real cost data.
- Quote: "Anthropic published the receipts, which almost nobody does."
- Quote: "A simplified version built a browser music app in 3 hours 50 minutes for $124.70. The planner cost 46 cents."
Playwright
- Description: Browser automation tool.
- Why mentioned: Used by Anthropic's evaluator agent to test the live app.
- Quote: "A third drives the finished app in a live browser through Playwright and grades it, catching bugs like wrong route ordering that would pass normal CI."
The AI Corner (publisher)
- Description: Newsletter selling a premium "harness manual."
- Why mentioned: Source of the article and its paid offering of templates and scorecards.
- Quote: "Premium subscribers get the full harness manual."
4. People Identified
Mitchell Hashimoto
- Description: Engineer who used the term "harness" on his blog.
- Why mentioned: Credited with helping popularize the term.
- Quote: "The term took off in February 2026, when Mitchell Hashimoto used it on his blog and OpenAI published its harness engineering post a few days later."
Prithvi Rajasekaran
- Description: Anthropic researcher who designed the three-agent harness.
- Why mentioned: Architect of the GAN-inspired role split.
- Quote: "Anthropic's Prithvi Rajasekaran borrowed the structure of a GAN. One agent, the planner, turns a one-sentence brief into a full product spec, and a second agent builds it."
Ruben Dominguez
- Description: Author byline for the newsletter post.
- Why mentioned: Writer of the analysis.
- Quote: "Ruben Dominguez"
5. Operating Insights
1. Audit the harness you already have
- Treat every retyped rule and every reflexive permission click as an undesigned harness component, and fix them deliberately.
- Quote: "Every rule you keep retyping into chat is part of your harness. So is every permission you clicked through because clicking was faster than reading."
2. Give agents a map, not a manual
- Keep the instruction file short (~100 lines) as an index into structured docs, and enforce architecture with linters rather than prose.
- Quote: "Give Codex a map, not a 1,000-page instruction manual."
3. Separate building from grading, and verify like a user
- Use an independent evaluator that exercises the real app, and have agents communicate via files for durable state. Gate permissions per job and log everything.
- Quote: "The three talk only by writing files to each other."
- Quote: "Every agent gets its own sandbox."
6. Overlooked Insights
1. One green run means little: reliability needs repeated evaluation
- The premium teaser references a "pass-cubed" eval method, implying agent reliability should be measured across repeated runs rather than single successes.
- Quote: "The pass-cubed eval method, and why one green run tells you almost nothing"
2. Planning is nearly free relative to the total cost
- In Anthropic's cheaper run, the planner cost 46 cents of a $124.70 total, suggesting specification is a very cheap place to add reliability.
- Quote: "The planner cost 46 cents."