Import AI 475: Swarm scaling; Google DeepMind watermarks biology; and the AI science economy
1. Key Themes
Theme: Agents are opening new scaling axes, with swarms trading tokens for speed
Swarms are a new form of inference-scaling that buys wall-clock time
Toby Ord's analysis reframes swarms as a scaling lever rather than a novelty. The tradeoff is more total tokens in exchange for faster completion.
"A good way to see AI swarms is as a new form of inference-scaling," he says.
"The 4-agent swarm needed about twice the total number of tokens to get the same performance, but in terms of tokens per agent, it only needed half as many. Since the agents are run in parallel, this means it can theoretically achieve the same task in half the time."
Swarms hit a "stepping on toes" coordination tax
Returns diminish as you add agents, mirroring human team dynamics, so 10x agents is not equal to 10x tokens on one agent.
"This means that scaling up the number of agents in the swarm by 10x doesn't get as much performance as using 10x as many tokens with one agent. Instead it gets 10λ x as much — which is 3x to 5x."
Clark sees upside if coordination improves:
"As we figure out how to get agents to productively coordinate... we could see even greater returns to scaling."
Theme: AI-driven science is moving from capability demos to physical-world operation
Frontier models can already pass a meaningful share of automated-lab tasks
SciUniverse tests models on 92 tasks across 17 families, from sample prep to instrument control to interpreting real measurements. Top performance is still under 50%, with substantial per-task costs.
"Claude Fable 5.1 (xhigh) leads with a pass rate of 45.3% and a cost-per-task of $40.61, followed by GPT-5 Astra (xhigh) at 32.5% and $52.37, then Claude Opus 5 (xhigh) at 30.5% and $46.31."
"I now think the same is true of how good AI systems are getting at performing increasingly automated science."
The bottleneck for AI science shifts to physical resources, motivating a "science economy"
DeepMind argues ideas will be abundant while lab capacity and empirical validation are scarce, so markets are needed to allocate them.
"The development of AI scientists is likely to be bottlenecked primarily by physical resources and empirical validation, rather than the ability to produce plausible or promising research ideas."
The proposed market has four components: proof of ideation, ex-ante evaluation (agents stake compute credits to forecast an idea's viability), brokerage and trade via fractional licensing, and validation payouts as automatic royalties.
"It would also make it possible to financially decouple the computational labour of ideation from the capital-intensive labour of physical execution."
Theme: Biosecurity and governance pressure is building around AI
Watermarking is emerging as one layer of AI-bio defense
Google DeepMind's SynthID Bio watermarks protein sequences and predicted 3D structures without apparent loss of function.
"In wet-lab testing across three target proteins (VEGF-A, the SARS-CoV-2 spike protein RBD, and PD-L1), our watermarked designs matched the hit rate, binding affinity, and natural sequence diversity of unwatermarked versions."
Clark frames it as one piece of a defense-in-depth stack:
"The SynthID Bio approach represents one thing to do here and will need to coordinate with broader monitoring of the physical equipment used to manufacture things, as well as AI-provider classifiers and other methods for reducing misuse."
Public opinion is running ahead of Washington on AI regulation
"61% of Americans (Sample size: 2498) think this agreement is 'not enough', including 53% of Trump voters."
"54% of voters (61% Harris, 48% Trump) said they thought 'the government should set and enforce rules for AI'."
2. Contrarian Perspectives
Voluntary industry commitments are viewed as insufficient, even by the president's own voters
Against the administration's approach of leaning on voluntary commitments from leading labs, polling shows majorities want enforceable rules, including a majority of Trump voters. Clark sees this as a latent political risk for the industry.
"This asymmetry between voter preferences and political action is inherently unstable. Energy is building up in the political system - what happens when it boils over?"
Swarms may accelerate, not restrain, an intelligence explosion
Ord had hoped that coordination costs among AI agents would be high enough to make recursive self-improvement less likely. The data point the other way.
"I'd hoped that the value of λ for AI agents would be lower, making an intelligence explosion less likely, but that appears to not be the case."
Despite diminishing returns, the swarm tax is mild enough (3x to 5x from 10x agents) that parallel scaling remains powerful.
Idea generation, not execution, is the cheap part of science
Conventional framing treats AI's value in science as producing better hypotheses. DeepMind argues the scarce input will be lab capacity and validation, so value shifts to those who control physical execution.
"...rather than the ability to produce plausible or promising research ideas."
3. Companies Identified
Google DeepMind
- Description: Google's AI research lab.
- Why mentioned: Released SynthID Bio for watermarking AI-designed biology and published a paper proposing an automated scientific economy.
- Quotes: "Google has developed SynthID Bio, a 'family of watermarking methods developed specifically for synthetic biology to strengthen biosecurity and scientific integrity'." / "The scientific community must proactively develop a native Automated Scientific Economy."
C5R Corp
- Description: Organization behind the SciUniverse benchmark.
- Why mentioned: Built a benchmark testing AI systems on operating a mostly automated scientific lab.
- Quotes: "These are some of the questions C5R Corp is trying to answer with SciUniverse, a way of testing out how well AI systems can operate a mostly automated scientific lab."
Anthropic
- Description: Frontier AI lab.
- Why mentioned: Its models (Claude Fable 5.1, Claude Opus 5) lead and place third on SciUniverse; also named among companies in the voluntary commitments announcement.
- Quotes: "Claude Fable 5.1 (xhigh) leads with a pass rate of 45.3%..." / "...including Anthropic and OpenAI..."
OpenAI
- Description: Frontier AI lab.
- Why mentioned: GPT-5 Astra placed second on SciUniverse; a signatory to the voluntary commitments.
- Quotes: "...followed by GPT-5 Astra (xhigh) at 32.5% and $52.37..."
Center for Shared AI Prosperity (CSAIP)
- Description: Polling organization.
- Why mentioned: Published polling showing public skepticism of voluntary AI self-governance.
- Quotes: "Polling from the Center for Shared AI Prosperity (CSAIP) suggests that Americans believe it is 'not enough' for companies to agree to self-police themselves about AI development."
HuggingFace
- Description: Open-source AI platform.
- Why mentioned: Cited as an example of agents coordinating to do more than the sum of their parts.
- Quotes: "agents can coordinate to do greater-than-sum-of-parts stuff, like the HuggingFace hack"
4. People Identified
Toby Ord
- Description: Researcher and author of the "Swarm Scaling" post.
- Why mentioned: Provided the core analysis of swarms as inference-scaling and quantified the coordination penalty.
- Quotes: "Why would you ever use swarms? The most important answer is speed." / "I'd hoped that the value of λ for AI agents would be lower, making an intelligence explosion less likely, but that appears to not be the case."
5. Operating Insights
Use swarms when time-to-result matters more than token cost
Parallelization buys speed at the price of more total tokens. For latency-sensitive workflows, budget roughly 2x tokens for a 4-agent swarm to halve wall-clock time, but don't expect linear returns.
"The 4-agent swarm needed about twice the total number of tokens to get the same performance, but in terms of tokens per agent, it only needed half as many."
Budget for diminishing returns when scaling agent count
Plan around the coordination tax rather than assuming agents scale like tokens. At large scale-ups, a single deeper agent may beat a wider swarm per token spent.
"...this shortfall accumulates quickly for larger scaleups, with the swarm falling further and further behind."
Expect automated-lab benchmarks to be cost-significant and not yet reliable
With the best model at 45.3% pass rate and roughly $40 per task, automated science workflows still need human oversight and cost modeling.
"Claude Fable 5.1 (xhigh) leads with a pass rate of 45.3% and a cost-per-task of $40.61."
6. Overlooked Insights
Agents could trade ideas as financial instruments via forecasting markets and fractional licensing
The DeepMind proposal for agents to stake compute credits on an idea's viability, then license it fractionally and collect royalties on validation, sketches a new asset class and marketplace structure that could create business opportunities in brokerage, evaluation, and physical-execution services.
"...agents and institutions to negotiate priority access to limited physical resources depending on the estimated value of different research proposals."
Watermarking must be validated against function, not just detectability
SynthID Bio's key proof point is that watermarked designs showed no loss in lab performance, which is the practical barrier to adoption by legitimate biotech users.
"...our watermarked designs matched the hit rate, binding affinity, and natural sequence diversity of unwatermarked versions."