Anthropic CEO’s Plan to Pace the Frontier: 3 Key Strategies

Matilda
7 Min Read

Dario Amodei, CEO of Anthropic, has outlined a plan to pace the frontier of AI development. In a new blog post, he echoed growing calls for caution and proposed three broad strategies for slowing the rapid improvement of AI models.

Amodei said two things convinced him it is time for a more cautious approach: the OpenAI-HuggingFace hack, and the fact that AI has been advancing drastically faster in recent months, particularly with its growing ability to build the next generation of AI.

“We must slow the pace at which we improve the capabilities of AI models,” Amodei wrote. “Progress will still seem fast, and we must make wise use of the time we gain.”

His proposal comes amid an intense debate over AI safety and alignment. The discussion intensified this week after researcher Jacob Coxon wrote that he is resigning from Anthropic over concerns that leading AI companies are “gambling with our lives” while the people building the technology “earnestly believe it could kill us all by the end of the decade,” a claim repeated by others at Anthropic.

Amodei’s post did not explicitly mention Coxon’s resignation or his concerns.

1. Embedded Evaluators

The first step in Amodei’s plan to pace the frontier involves “embedded evaluators” from third-party organizations like METR. These evaluators would verify that AI companies are actually following their pacing and safety commitments and ensure that safety incidents get reported.

OpenAI was recently criticized for not reporting an incident where its AI agents took over a German wiki form.

Amodei compared these evaluators to regulators who have been embedded with bank employees. He said that inviting them in is “something Anthropic is unilaterally committing to (and calls on governments to require other frontier companies to match).”

That means giving evaluators company badges, desks, and laptops, and providing access “mostly comparable to what internal risk assessment teams have,” with exceptions when required by law or contracts.

2. Coordination Among Democratic Countries

Next, Amodei called for the leading AI companies “within democratic countries” to coordinate “common safety standards as well as limits on the rate of unchecked AI progress.”

Such coordination might seem unlikely, both due to the apparent animosity between OpenAI CEO Sam Altman and Amodei and because their companies are reportedly worried that a coordinated pause could lead to antitrust scrutiny.

Amodei alluded to that concern in his post, writing that “for antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions — they don’t need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations.”

Amodei also acknowledged the spectre of Chinese AI dominance that is often raised as an argument against slowing development. But he said that if the US government and tech companies take steps like refusing to sell powerful chips or semiconductor manufacturing equipment to Chinese companies, as well as cracking down on model distillation, they could “slow China’s progress enough to widen America’s lead significantly over the next 3–5 years.”

3. Global Coordination

Lastly, Amodei called for “global coordination,” where the United States and its allies “attempt to coordinate with authoritarian governments, to the extent this is possible.”

Amodei said this would mean “cooperation with China,” and he admitted that there are “stark limits on what can be achieved.” But he still suggested there might be opportunities for agreement, even if it is just “prohibiting certain narrow and obviously dangerous uses of AI, such as using AI for the production of biological weapons or allowing users to do so.”

A Crisis of Trust

With Amodei’s past willingness to acknowledge AI’s potential dangers, and with the company’s relative openness to certain forms of regulation, some AI boosters have already criticized him as a doomer whose comments have fed the current AI backlash.

In response, Amodei said he has tried to offer a “balanced” perspective and argued that the backlash is “fundamentally a crisis of trust,” as people have become skeptical of tech companies, the tech industry, and the government.

Industry critics have also been skeptical about these apocalyptic AI warnings, suggesting that they are a distraction from the harm that the technology is already causing.

Journalist Brian Merchant, for example, wrote that he has yet to see “a credible, step-by-step documentation of how exactly AI might move from self-recursively improving AI to killing every single human on the planet.” He also suggested that proposals similar to Amodei’s “would likely only wind up serving Anthropic and OpenAI; it’s what regulatory capture looks like in action.”

Amodei’s Stated Intentions

In his new post, Amodei wrote that he continues “to believe that AI can enormously improve the quality of human life.”

“My desire to achieve these benefits is undimmed,” he said. “But the benefits will only be achieved if we build the technology in the right way, and — so long as we use the time we gain well — it is worth taking unusually deliberate care to get it right.”

The debate over how to pace the frontier of AI development is far from settled. Amodei’s proposal represents one of the most detailed roadmaps yet from a leading AI CEO. Whether it gains traction among other companies and governments remains to be seen.

Share This Article
Leave a Comment

Leave a Reply

Your email address will not be published. Required fields are marked *