Within Safety Thresholds

Could Shared Rules Stop an AI Safety Race?

Common rules, external testing and regulation could reduce the advantage gained by weakening safety thresholds during an AI race.

42 sources 3 graphics
Preview for Could Shared Rules Stop an AI Safety Race?

On this page

  • Why unilateral restraint can punish cautious developers
  • How common evaluations could reduce competitive pressure
  • Where regulation and international coordination may fail

Introduction

If capability thresholds are meant to slow or stop the deployment of unusually dangerous AI systems, they work best when they apply to everyone who matters. A single company that delays a powerful model for extra safety testing may lose customers, investment or strategic influence if competitors continue without comparable constraints. Shared rules aim to change those incentives. Instead of rewarding the fastest developer regardless of safety practices, they try to ensure that every major frontier developer faces similar obligations once models reach agreed capability thresholds.

Shared Standards illustration 1
Explanatory illustration 1

Within debates about AI doom and existential risk, this idea is attractive because many worrying scenarios depend on intense competition undermining caution. The hope is not that international agreements eliminate rivalry, but that common evaluations, external oversight and coordinated governance reduce the advantage gained by weakening safety standards. Whether this can succeed remains one of the central disputes in frontier AI governance.

Why unilateral restraint can punish cautious developers

Capability thresholds only influence behaviour if organisations are willing to accept the costs of complying with them. Those costs can include delayed product launches, additional security work, independent evaluations or, in exceptional cases, pausing development until risks are better understood.

The problem is that these costs are usually borne by one organisation, while many of the benefits are shared across society. Economists describe this as a coordination problem. Even if every laboratory would prefer all competitors to adopt stronger safeguards, each individual laboratory has an incentive to worry that slowing down alone simply hands an advantage to rivals.

This tension is recognised explicitly in several frontier AI policies. Anthropic originally argued that responsible scaling could encourage a “race to the top” if widely adopted, rather than rewarding the fastest deployment regardless of safety. Its Responsible Scaling Policy was presented as a framework that could eventually become an industry norm rather than a company-specific promise.[anthropic.com]anthropic.comAnnouncing Anthropic's Responsible Scaling Policy \ AnthropicSeptember 19, 2023…Published: September 19, 2023

OpenAI’s updated Preparedness Framework acknowledges the same competitive pressure from another direction. It states that if another frontier developer deploys a similarly risky system without equivalent safeguards, OpenAI may adjust some requirements, although only after publicly explaining the decision and concluding that overall severe-harm risk has not materially increased. The inclusion of such a clause illustrates how difficult purely unilateral commitments can become during an active technological race.[OpenAI]OpenAIupdating our preparedness frameworkOur updated Preparedness Framework | OpenAIApril 15, 2025…Published: April 15, 2025

For people concerned about AI doom, this is significant because existential-risk arguments often assume that incentives become strongest precisely when systems begin approaching genuinely dangerous capabilities.

How common evaluations could reduce competitive pressure

Shared capability thresholds do not require every organisation to build identical systems or follow identical internal processes. Instead, the goal is to create enough consistency that developers cannot easily gain an advantage simply by redefining risk.

Several mechanisms are commonly proposed.

Common capability evaluations. If laboratories measure dangerous capabilities using comparable testing methods, disagreements become easier to examine publicly. Rather than each company privately deciding whether its own model has crossed an important threshold, independent researchers and regulators can compare results across developers.

Independent external testing. Government-backed AI Safety Institutes and accredited third-party evaluators can provide additional scrutiny. External assessments reduce the perception that companies are marking their own homework, while helping establish more consistent expectations for what counts as crossing a threshold.[arXiv]arxiv.orgThe Role of AI Safety Institutes in Contributing to International Standards for Frontier AI SafetySeptember 17, 2024…Published: September 17, 2024

Comparable reporting requirements. Publishing capability evaluations, risk reports and deployment decisions makes it harder for one developer to quietly weaken standards while competitors continue investing in safety. Transparency also allows researchers to identify whether thresholds are drifting over time.

Shared terminology. If developers broadly agree on what counts as dangerous cyber capability, biological assistance, autonomous AI research or loss-of-control indicators, disputes become focused on evidence rather than definitions.

None of these measures guarantees compliance. However, they can reduce opportunities for strategic reinterpretation and increase the reputational cost of abandoning agreed standards.

1:00:57

Why regulation may matter more than voluntary promises

Many AI safety researchers argue that voluntary commitments alone are unlikely to remain stable once commercial or geopolitical competition intensifies.

This does not necessarily imply bad faith. Companies answer to investors, employees, governments and customers while competing in rapidly changing markets. Even organisations with strong safety cultures may struggle to justify lengthy delays if competitors face no equivalent obligations.

Regulation attempts to alter these incentives by making safety requirements apply across an entire market instead of to individual firms.

Potential approaches include:

  • mandatory reporting when predefined capability thresholds are reached;
  • minimum evaluation requirements before deployment;
  • independent audits of frontier systems;
  • legal obligations to maintain adequate security for model weights and infrastructure;
  • penalties for ignoring agreed testing or disclosure requirements.

The underlying logic resembles safety regulation in other high-risk industries. Aviation manufacturers do not generally compete by eliminating aircraft certification, because certification requirements apply broadly across the industry.

Advocates of stronger governance argue that frontier AI may eventually require similar baseline expectations if capability thresholds are to remain credible under competitive pressure. The Seoul AI Safety Summit reflected this thinking by encouraging companies to publish frontier safety frameworks and commit not to deploy models whose severe risks could not be adequately mitigated, while governments expanded cooperation between national AI Safety Institutes.[Reuters]reuters.comAI summit secures safety commitments from 16 companiesThis announcement came during a global AI summit co-hosted by South Korea and Britain in Seoul. The commitment entails publishing safety…

Shared Standards illustration 2
Explanatory illustration 2

International coordination is harder than national regulation

National regulation alone may not solve the problem if frontier development simply shifts to jurisdictions with weaker requirements.

This creates another coordination challenge. A country imposing strict capability thresholds could fear losing investment, talent or strategic influence if rival countries allow unrestricted deployment.

As a result, proposals often focus on partial international coordination rather than perfectly uniform global regulation.

Possible forms include:

  • mutually recognised evaluation standards;
  • shared technical benchmarks for dangerous capabilities;
  • information-sharing between AI Safety Institutes;
  • coordinated incident reporting;
  • agreements on minimum security practices for frontier models;
  • interoperability between national regulatory systems.

These arrangements would not eliminate competition. Instead, they seek to reduce incentives for regulatory arbitrage, where developers choose locations primarily because safety requirements are weaker.

2:31:00

Where coordination could still fail

Shared standards are not a complete answer to competitive pressure.

Countries may disagree about acceptable risk

Governments have different economic priorities, security concerns and attitudes towards technological leadership.

A country that believes advanced AI offers decisive military or economic advantages may be reluctant to accept international limits that appear to slow domestic developers.

This becomes especially difficult if policymakers disagree about the probability of existential risk itself. Those assigning very low probabilities to AI doom may view strict thresholds as imposing large economic costs for uncertain benefits.

Dangerous capabilities remain difficult to measure

Even if governments agree in principle, technical disagreements remain.

Questions such as whether a model genuinely automates advanced cyber operations or meaningfully accelerates AI research are rarely answered by a single benchmark. Capability often depends on tools, prompting, fine-tuning and human collaboration.

Consequently, shared thresholds require ongoing technical work rather than one permanent numerical standard.

Shared Standards illustration 3
Explanatory illustration 3

Voluntary coordination may weaken over time

Recent developments illustrate this difficulty.

Anthropic’s original Responsible Scaling Policy explicitly discussed temporary pauses if safety procedures failed to keep pace with capabilities and argued that broad adoption could create better competitive incentives. As competitive conditions evolved, later revisions shifted towards greater transparency, published risk reports and governance documentation while removing some earlier unilateral commitments. Company leaders argued that isolated restraint becomes increasingly difficult without wider regulatory coordination, whereas critics saw the changes as evidence that voluntary promises erode under competitive pressure.[anthropic.com]anthropic.comAnnouncing Anthropic's Responsible Scaling Policy \ AnthropicSeptember 19, 2023…Published: September 19, 2023

For supporters of stronger governance, this episode demonstrates the central argument for shared rules: if even organisations identified with AI safety struggle to maintain unilateral commitments, durable safeguards may require broader institutional support.

What this means for AI doom debates

Supporters of shared capability thresholds do not generally argue that common rules eliminate existential risk. Instead, they argue that they improve the odds that safety decisions survive competitive pressure.

The hoped-for effect is a shift from competing over who deploys first to competing over who can satisfy increasingly demanding safety requirements. In that vision, external evaluations, common benchmarks and regulatory oversight transform safety from a commercial disadvantage into part of the competitive landscape.

Critics remain sceptical. They question whether governments can agree quickly enough, whether technical evaluations can keep pace with rapidly improving models, and whether geopolitical rivalry will ultimately outweigh cooperative governance.

These uncertainties explain why shared standards occupy an important place in discussions of AI doom. The debate is not simply about writing better capability thresholds, but about whether any developer can realistically keep them intact if competitors are free to ignore them. The stronger and more widely shared the rules become, the less likely it is that the next major advance will reward whichever organisation is most willing to lower its safety bar.

Amazon book picks

Further Reading

Books and field guides related to Could Shared Rules Stop an AI Safety Race?. Use these as the next step if you want deeper reading beyond the article.

BookCover for Human Compatible

Human Compatible

By Stuart Russell

A leading artificial intelligence researcher lays out a new approach to AI that will enable us to coexist successfully with increasingly...

BookCover for The Age of A. I.

The Age of A. I.

By Henry Kissinger, Eric Schmidt et al.

An A.I. that learned to play chess discovered moves that no human champion would have conceived of. Driverlesscars edge forward at red li...

BookCover for The Coming Wave

The Coming Wave

By Mustafa Suleyman

"We are approaching a critical threshold in the history of our species. Everything is about to change. Soon you will live surrounded by A...

BookCover for The Precipice

The Precipice

By Toby Ord

What existential threats does humanity face? And how can we secure our future?'The Precipice is a powerful book . . . Ord's love for huma...

eBay marketplace picks

Marketplace Samples

Live-tested eBay searches with available results related to this page.

UsingUSA

Selected fromrobotics wall art oneBay.co.uk.

Endnotes

1. Source: anthropic.com
Title: Announcing Anthropic’s Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/news/anthropics-responsible-scaling-policy

Source snippet

September 19, 2023...

Published: September 19, 2023

2. Source: anthropic.com
Title: ’s Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/responsible-scaling-policy

Source snippet

Anthropic’s Responsible Scaling Policy \ AnthropicJuly 8, 2026 — ANTHROPIC’S RESPONSIBLE SCALING POLICY Anticipating and securing against...

Published: July 8, 2026

3. Source: OpenAI
Title: updating our preparedness framework
Link:https://openai.com/index/updating-our-preparedness-framework/

Source snippet

Our updated Preparedness Framework | OpenAIApril 15, 2025...

Published: April 15, 2025

4. Source: arxiv.org
Link:https://arxiv.org/abs/2409.11314

Source snippet

The Role of AI Safety Institutes in Contributing to International Standards for Frontier AI SafetySeptember 17, 2024...

Published: September 17, 2024

5. Source: reuters.com
Title: AI summit secures safety commitments from 16 companies
Link:https://www.reuters.com/technology/global-ai-summit-seoul-aims-forge-new-regulatory-agreements-2024-05-21/

Source snippet

This announcement came during a global AI summit co-hosted by South Korea and Britain in Seoul. The commitment entails publishing safety...

6. Source: time.com
Title: exclusive anthropic drops flagship safety pledge
Link:https://time.com/7380854/exclusive-anthropic-drops-flagship-safety-pledge/

Source snippet

This pledge had promised to halt training of AI models unless safety measures could be ensured in advance. The company now believes such...

7. Source: OpenAI
Title: helping build shared standards for advanced ai
Link:https://openai.com/index/helping-build-shared-standards-for-advanced-ai/

Source snippet

comHelping build shared standards for advanced AI | OpenAIJune 23, 2026 — June 23, 2026 Global Affairs HELPING BUILD SHARED STANDARDS FOR...

Published: June 23, 2026

8. Source: OpenAI
Title: frontier governance framework
Link:https://openai.com/index/openai-frontier-governance-framework/

9. Source: anthropic.com
Title: The case for targeted regulation \ Anthropic
Link:https://www.anthropic.com/news/the-case-for-targeted-regulation

10. Source: anthropic.com
Title: Reflections on our Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/news/reflections-on-our-responsible-scaling-policy

11. Source: OpenAI
Title: our approach to frontier risk
Link:https://openai.com/global-affairs/our-approach-to-frontier-risk/

12. Source: OpenAI
Link:https://openai.com/safety/how-we-think-about-safety-alignment/

Additional References

13. Source: apnews.com
Link:https://apnews.com/article/2cc2b297872d860edc60545d5a5cf598

Source snippet

This gathering is a continuation of the AI Safety Summit held in the UK in November, reflecting global efforts to create safeguards again...

14. Source: youtube.com
Title: Are Anthropic’s AI safety policies up to the task? | Nick Joseph
Link:https://www.youtube.com/watch?v=E6_x0ZOXVVI

Source snippet

Claude Mythos can make the world 'much more secure' says Anthropic co-founder...

15. Source: oecd.org
Link:https://www.oecd.org/en/topics/ai-principles.html

16. Source: youtube.com
Title: Claude Mythos can make the world ‘much more secure’ says Anthropic co-founder
Link:https://www.youtube.com/watch?v=7JYb0JGz7rg

Source snippet

Paul Christiano — Preventing an AI takeover...

17. Source: youtube.com
Title: Demis Hassabis Is Right About AI Risk—But Testing Is Not Enough
Link:https://www.youtube.com/watch?v=I0qhT3es5GY

Source snippet

Adaptive AI Governance with Gillian Hadfield and Andrew Freedman...

18. Source: nist.gov
Link:https://www.nist.gov/news-events/news/2024/11/fact-sheet-us-department-commerce-us-department-state-launch-international

19. Source: nist.gov
Link:https://www.nist.gov/news-events/news/2024/05/us-secretary-commerce-gina-raimondo-releases-strategic-vision-ai-safety

20. Source: youtube.com
Title: Adaptive AI Governance with Gillian Hadfield and Andrew Freedman
Link:https://www.youtube.com/watch?v=s7jAqAsu1t4

Source snippet

Can we escape Moloch's trap with a GPU treaty?...

21. Source: oecd.org
Title: enablers guardrails and engagement for unlocking trustworthy ai 2f817983
Link:https://www.oecd.org/en/publications/2025/06/governing-with-artificial-[intelligence

22. Source: oecd.org
Title: enablers guardrails and engagement for unlocking trustworthy ai 2f817983
Link:https://www.oecd.org/en/publications/governing-with-artificial-intelligence_795de142-en/full-report/enablers-guardrails-and-engagement-for-unlocking-trustworthy-ai_2f817983.html