Within AI Race

Will AI Safety Thresholds Survive Competitive Pressure?

Safety thresholds only slow dangerous development if laboratories apply them consistently when pausing could cost a competitive lead.

23 sources 3 graphics
Preview for Will AI Safety Thresholds Survive Competitive Pressure?

On this page

  • What capability thresholds are meant to stop
  • Where internal judgement and disclosure can fail
  • How shared standards could reduce pressure to defect

Introduction

Capability thresholds are one of the main ideas behind modern frontier AI safety policies. Rather than requiring every powerful model to meet the same safeguards, they specify that once a system demonstrates particular dangerous capabilities—such as substantially improving cyber attacks or automating advanced AI research—it should trigger stronger testing, tighter security, or even delayed deployment until additional protections are in place. The central question is not whether thresholds can be written down, but whether organisations will still honour them when delaying a release could mean losing commercial, geopolitical or strategic advantage.

Safety Thresholds illustration 1

Within AI doom discussions, this question matters because many catastrophic-risk scenarios assume that warning signs appear gradually rather than all at once. If laboratories reliably pause when models cross agreed capability thresholds, there may be time to strengthen oversight before more dangerous systems are widely deployed. If competitive pressure encourages firms to reinterpret, weaken or ignore their own thresholds, then the practical value of those safeguards could be far lower than their public commitments suggest.

What capability thresholds are meant to stop

Capability thresholds are intended to convert vague promises about “being careful” into operational decisions. Instead of asking whether a model is generally powerful, they ask whether it can perform specific tasks associated with catastrophic risks.

Several frontier laboratories have adopted versions of this approach:

  • Anthropic’s Responsible Scaling Policy defines capability thresholds linked to risks such as advanced cyber capabilities, chemical and biological misuse, and increasingly autonomous AI research. Crossing a threshold requires stronger security standards, additional evaluations and governance procedures before deployment.[anthropic.com]anthropic.com’s Responsible Scaling Policy \ AnthropicAnthropic’s Responsible Scaling Policy \ AnthropicJuly 8, 2026…Published: July 8, 2026
  • OpenAI’s Preparedness Framework similarly classifies high-risk capabilities and states that models reaching defined “High” or “Critical” capability levels should satisfy increasingly demanding safety requirements before deployment or, for the highest-risk systems, during development itself.[OpenAI]OpenAIupdating our preparedness frameworkOur updated Preparedness Framework | OpenAIApril 15, 2025…Published: April 15, 2025

From an AI doom perspective, the most important thresholds are not ordinary product improvements but capabilities that could rapidly increase existential risk, for example:

  • substantially accelerating AI research itself;
  • enabling sophisticated cyber operations at scale;
  • making catastrophic biological misuse substantially easier;
  • demonstrating behaviours suggesting persistent loss of human control.

The logic is deliberately preventative. Waiting until obvious harm occurs may be too late if the capabilities involved improve rapidly.

Why competitive pressure makes thresholds difficult to enforce

The challenge is not writing a threshold but acting on it.

Suppose two frontier laboratories are developing similarly capable models. If one concludes its newest model has crossed an internal safety threshold requiring weeks or months of additional testing, that delay may allow its competitor to release first, attract customers, recruit talent and shape industry standards.

Even if both companies genuinely prefer stronger safety, each may worry that unilateral restraint simply benefits rivals. Economists describe this as a coordination problem rather than necessarily bad faith.

The incentives become even stronger if organisations believe that:

  • market leadership produces lasting competitive advantages;
  • governments are rewarding rapid capability growth;
  • investors expect continual product releases;
  • geopolitical competition makes slowing down appear strategically costly.

In this environment, thresholds can become increasingly difficult to apply consistently precisely when they matter most.

7:51

Where internal judgement and disclosure can fail

Capability thresholds are often presented as objective trigger points, but in practice they involve considerable judgement.

Measuring the capability is not straightforward

Many dangerous abilities cannot be measured as simply as benchmark scores.

Questions may include:

  • Does a model merely assist advanced cyber attacks, or can it independently conduct them?
  • Is improved AI coding enough to count as automating AI research?
  • How much tool use or scaffolding should count when evaluating dangerous capability?

As systems improve gradually rather than suddenly, deciding exactly when a threshold has been crossed becomes increasingly difficult.

Anthropic has publicly acknowledged this problem, noting that determining whether certain AI research capability thresholds have been crossed is becoming more subjective as frontier models improve.[anthropic.com]anthropic.com’s Responsible Scaling Policy \ AnthropicAnthropic’s Responsible Scaling Policy \ AnthropicJuly 8, 2026…Published: July 8, 2026

Safety Thresholds illustration 2

The laboratory often judges its own model

Another concern is institutional independence.

Most current frontier safety frameworks rely heavily on internal evaluations performed by the same organisation that wishes to deploy the model. External experts may participate in reviews, but the developer frequently controls:

  • which evaluations are performed;
  • what evidence is disclosed publicly;
  • how uncertain results are interpreted;
  • whether deployment proceeds.

This does not mean companies act dishonestly. However, governance scholars have pointed out that self-assessment naturally creates incentives to interpret ambiguous evidence in ways that permit deployment rather than delay it. Independent evaluations remain relatively limited across the industry.[arXiv]arxiv.orgEvaluating AI Companies' Frontier Safety Frameworks: Methodology and ResultsDecember 1, 2025…Published: December 1, 2025

Public transparency may be incomplete

Many safety commitments depend on information outsiders cannot easily verify.

External observers rarely know:

  • the complete evaluation results;
  • the exact internal debate before deployment;
  • capabilities discovered after release;
  • whether commercial considerations influenced decisions.

This makes it difficult for regulators, researchers or the public to assess whether capability thresholds are consistently applied.

7:41

Recent changes illustrate the pressure

The evolution of frontier safety policies has itself become evidence in this debate.[anthropic.com]anthropic.comFrontier Safety Roadmap \ AnthropicFrontier Safety Roadmap \ Anthropic

Anthropic originally framed its Responsible Scaling Policy around strong commitments not to proceed when required safeguards were unavailable. Later revisions retained extensive capability thresholds and transparency measures but removed some earlier language that critics interpreted as implying automatic pauses. Company leaders argued that unilateral commitments had become increasingly impractical without broader coordination and that transparency combined with evolving safeguards offered a more realistic governance approach.[anthropic.com]anthropic.com’s Responsible Scaling Policy \ AnthropicAnthropic’s Responsible Scaling Policy \ AnthropicJuly 8, 2026…Published: July 8, 2026

Supporters argue this reflects realism rather than retreat. They note that a single cautious laboratory cannot safely slow development if competitors with weaker standards continue advancing unchecked.

Critics respond that this illustrates exactly the concern capability thresholds were meant to solve: competitive pressure can gradually weaken voluntary commitments even among organisations regarded as unusually safety-focused.

The episode does not demonstrate that thresholds inevitably fail, but it highlights how difficult maintaining stringent commitments becomes during an active technological race.

How shared standards could reduce pressure to defect

Many researchers argue that capability thresholds work best when they are not unique to one laboratory.

If competing organisations use similar definitions, evaluation methods and required safeguards, the commercial penalty for acting cautiously becomes smaller.

Possible approaches include:

  • common definitions of high-risk capabilities across major developers;
  • shared evaluation protocols so thresholds are measured consistently;
  • independent external assessment of frontier models;
  • regulatory requirements that apply similar obligations to all major developers;
  • coordinated international agreements covering frontier systems rather than relying entirely on voluntary promises.

Recent research has argued that inconsistent capability thresholds across companies make comparison difficult and may encourage a “race to the bottom”, proposing more harmonised approaches for defining safety triggers.[arXiv]arxiv.orgarXiv Harmonizing AI Safety ThresholdsHarmonizing AI Safety ThresholdsJuly 17, 2026…Published: July 17, 2026

The objective is not necessarily identical policies, but enough consistency that no developer gains a major competitive advantage simply by applying weaker standards.

Safety Thresholds illustration 3

Can capability thresholds actually survive an AI race?

The evidence so far does not justify either optimism or pessimism alone.

Capability thresholds have become significantly more sophisticated than early voluntary AI principles. Leading laboratories increasingly publish concrete governance frameworks, specify categories of dangerous capability and describe corresponding safety measures. This represents genuine progress compared with broad statements of intent.[OpenAI]OpenAIupdating our preparedness frameworkOur updated Preparedness Framework | OpenAIApril 15, 2025…Published: April 15, 2025

At the same time, no existing framework has been tested under the conditions that concern AI doom researchers most: a model that clearly offers transformative strategic advantage while simultaneously triggering unprecedented safety concerns.

Whether thresholds survive that moment depends less on the wording of the policies than on the incentives surrounding them.

If major competitors expect others to respect comparable standards, capability thresholds could become practical coordination tools that buy valuable time for evaluation and stronger safeguards. If commercial or geopolitical competition rewards whichever organisation moves first, even carefully designed thresholds may gradually be reinterpreted, delayed or weakened.

For those concerned about existential AI risk, this is why capability thresholds are viewed as necessary but not sufficient. They provide a mechanism for translating dangerous capabilities into concrete safety actions, but their effectiveness ultimately depends on whether competitive pressures allow organisations to honour those commitments when the costs of caution become greatest.

Amazon book picks

Further Reading

Books and field guides related to Will AI Safety Thresholds Survive Competitive Pressure?. Use these as the next step if you want deeper reading beyond the article.

BookCover for Human Compatible

Human Compatible

By Stuart Russell

A leading artificial intelligence researcher lays out a new approach to AI that will enable us to coexist successfully with increasingly...

BookCover for The Alignment Problem

The Alignment Problem

By Brian Christian

Finalist for the Los Angeles Times Book Prize A jaw-dropping exploration of everything that goes wrong when we build AI systems and the m...

BookCover for The Coming Wave

The Coming Wave

By Mustafa Suleyman

"We are approaching a critical threshold in the history of our species. Everything is about to change. Soon you will live surrounded by A...

BookCover for The Black Box Society

The Black Box Society

By Frank Pasquale

Every day, corporations are connecting the dots about our personal behavior—silently scrutinizing clues left behind by our work habits an...

eBay marketplace picks

Marketplace Samples

Live-tested eBay searches with available results related to this page.

UsingUSA

Selected fromcybersecurity patch oneBay.co.uk.

Endnotes

1. Source: anthropic.com
Title: ’s Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/responsible-scaling-policy

Source snippet

Anthropic’s Responsible Scaling Policy \ AnthropicJuly 8, 2026...

Published: July 8, 2026

2. Source: anthropic.com
Title: Announcing our updated Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/news/announcing-our-updated-responsible-scaling-policy

3. Source: OpenAI
Title: updating our preparedness framework
Link:https://openai.com/index/updating-our-preparedness-framework/

Source snippet

Our updated Preparedness Framework | OpenAIApril 15, 2025...

Published: April 15, 2025

4. Source: OpenAI
Title: Open AIOpen AI’s Frontier Governance Framework | Open AI
Link:https://openai.com/index/openai-frontier-governance-framework/

Source snippet

’s Frontier Governance Framework | OpenAI...

5. Source: arxiv.org
Link:https://arxiv.org/abs/2512.01166

Source snippet

Evaluating AI Companies' Frontier Safety Frameworks: Methodology and ResultsDecember 1, 2025...

Published: December 1, 2025

6. Source: arxiv.org
Link:https://arxiv.org/abs/2509.24394

Source snippet

The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analys...

7. Source: time.com
Title: exclusive anthropic drops flagship safety pledge
Link:https://time.com/7380854/exclusive-anthropic-drops-flagship-safety-pledge/

Source snippet

This pledge had promised to halt training of AI models unless safety measures could be ensured in advance. The company now believes such...

8. Source: arxiv.org
Title: arXiv Harmonizing AI Safety Thresholds
Link:https://arxiv.org/abs/2607.16112

Source snippet

Harmonizing AI Safety ThresholdsJuly 17, 2026...

Published: July 17, 2026

9. Source: arxiv.org
Title: arXiv Risk thresholds for frontier AI
Link:https://arxiv.org/abs/2406.14713

10. Source: anthropic.com
Title: Frontier Safety Roadmap \ Anthropic
Link:https://www.anthropic.com/responsible-scaling-policy/roadmap

11. Source: OpenAI
Title: frontier safety blueprint
Link:https://openai.com/index/frontier-safety-blueprint/

12. Source: anthropic.com
Title: Reflections on our Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/news/reflections-on-our-responsible-scaling-policy

13. Source: OpenAI
Title: frontier risk and preparedness
Link:https://openai.com/index/frontier-risk-and-preparedness/

14. Source: OpenAI
Title: our approach to frontier risk
Link:https://openai.com/global-affairs/our-approach-to-frontier-risk/

15. Source: youtube.com
Title: Anthropic’s AI Safety Plan
Link:https://www.youtube.com/watch?v=Z_nHHKrcjQM

Source snippet

OpenAI's Preparedness Framework: AI Safety Plan...

16. Source: youtube.com
Title: Open AI’s Preparedness Framework: AI Safety Plan
Link:https://www.youtube.com/watch?v=Mx07W9M60Gs

Source snippet

The Catastrophic Risks of AI — and a Safer Path...

Additional References

17. Source: businessinsider.com
Link:https://www.businessinsider.com/anthropic-changing-safety-policy

Source snippet

CEO Dario Amodei emphasized that the original policy was never intended to be static and noted practical challenges in enforcing high-ris...

18. Source: youtube.com
Title: The Catastrophic Risks of AI — and a Safer Path
Link:https://www.youtube.com/watch?v=qe9QSCF-d88

Source snippet

MEPs & AI expert reveal what it would actually take to pause AI development...

19. Source: deepmind.google
Title: Google Deep Mind strengthens the Frontier Safety Framework — Google Deep Mind
Link:https://deepmind.google/blog/strengthening-our-frontier-safety-framework/

20. Source: youtube.com
Title: Establishing AI Risk Thresholds: A Critical Global Policy Priority
Link:https://www.youtube.com/watch?v=ThN7GY0CjHY

Source snippet

Anthropic's AI Safety Plan...

21. Source: deepmind.google
Title: Introducing the Frontier Safety Framework — Google Deep Mind
Link:https://deepmind.google/blog/introducing-the-frontier-safety-framework/

22. Source: youtube.com
Title: MEPs & AI expert reveal what it would actually take to pause AI development
Link:https://www.youtube.com/watch?v=gCM5mrJOYp8

23. Source: openropic.com
Title: Responsible Scaling Policy Updates
Link:https://openropic.com/responsible-scaling-policy