Within AI Race
Will AI Safety Thresholds Survive Competitive Pressure?
Safety thresholds only slow dangerous development if laboratories apply them consistently when pausing could cost a competitive lead.
On this page
- What capability thresholds are meant to stop
- Where internal judgement and disclosure can fail
- How shared standards could reduce pressure to defect
Page outline Jump by section
Introduction
Capability thresholds are one of the main ideas behind modern frontier AI safety policies. Rather than requiring every powerful model to meet the same safeguards, they specify that once a system demonstrates particular dangerous capabilities—such as substantially improving cyber attacks or automating advanced AI research—it should trigger stronger testing, tighter security, or even delayed deployment until additional protections are in place. The central question is not whether thresholds can be written down, but whether organisations will still honour them when delaying a release could mean losing commercial, geopolitical or strategic advantage.
Within AI doom discussions, this question matters because many catastrophic-risk scenarios assume that warning signs appear gradually rather than all at once. If laboratories reliably pause when models cross agreed capability thresholds, there may be time to strengthen oversight before more dangerous systems are widely deployed. If competitive pressure encourages firms to reinterpret, weaken or ignore their own thresholds, then the practical value of those safeguards could be far lower than their public commitments suggest.
What capability thresholds are meant to stop
Capability thresholds are intended to convert vague promises about “being careful” into operational decisions. Instead of asking whether a model is generally powerful, they ask whether it can perform specific tasks associated with catastrophic risks.
Several frontier laboratories have adopted versions of this approach:
- Anthropic’s Responsible Scaling Policy defines capability thresholds linked to risks such as advanced cyber capabilities, chemical and biological misuse, and increasingly autonomous AI research. Crossing a threshold requires stronger security standards, additional evaluations and governance procedures before deployment.[anthropic.com]anthropic.com’s Responsible Scaling Policy \ AnthropicAnthropic’s Responsible Scaling Policy \ AnthropicJuly 8, 2026…
- OpenAI’s Preparedness Framework similarly classifies high-risk capabilities and states that models reaching defined “High” or “Critical” capability levels should satisfy increasingly demanding safety requirements before deployment or, for the highest-risk systems, during development itself.[OpenAI]OpenAIupdating our preparedness frameworkOur updated Preparedness Framework | OpenAIApril 15, 2025…
From an AI doom perspective, the most important thresholds are not ordinary product improvements but capabilities that could rapidly increase existential risk, for example:
- substantially accelerating AI research itself;
- enabling sophisticated cyber operations at scale;
- making catastrophic biological misuse substantially easier;
- demonstrating behaviours suggesting persistent loss of human control.
The logic is deliberately preventative. Waiting until obvious harm occurs may be too late if the capabilities involved improve rapidly.
Why competitive pressure makes thresholds difficult to enforce
The challenge is not writing a threshold but acting on it.
Suppose two frontier laboratories are developing similarly capable models. If one concludes its newest model has crossed an internal safety threshold requiring weeks or months of additional testing, that delay may allow its competitor to release first, attract customers, recruit talent and shape industry standards.
Even if both companies genuinely prefer stronger safety, each may worry that unilateral restraint simply benefits rivals. Economists describe this as a coordination problem rather than necessarily bad faith.
The incentives become even stronger if organisations believe that:
- market leadership produces lasting competitive advantages;
- governments are rewarding rapid capability growth;
- investors expect continual product releases;
- geopolitical competition makes slowing down appear strategically costly.
In this environment, thresholds can become increasingly difficult to apply consistently precisely when they matter most.
Where internal judgement and disclosure can fail
Capability thresholds are often presented as objective trigger points, but in practice they involve considerable judgement.
Measuring the capability is not straightforward
Many dangerous abilities cannot be measured as simply as benchmark scores.
Questions may include:
- Does a model merely assist advanced cyber attacks, or can it independently conduct them?
- Is improved AI coding enough to count as automating AI research?
- How much tool use or scaffolding should count when evaluating dangerous capability?
As systems improve gradually rather than suddenly, deciding exactly when a threshold has been crossed becomes increasingly difficult.
Anthropic has publicly acknowledged this problem, noting that determining whether certain AI research capability thresholds have been crossed is becoming more subjective as frontier models improve.[anthropic.com]anthropic.com’s Responsible Scaling Policy \ AnthropicAnthropic’s Responsible Scaling Policy \ AnthropicJuly 8, 2026…
The laboratory often judges its own model
Another concern is institutional independence.
Most current frontier safety frameworks rely heavily on internal evaluations performed by the same organisation that wishes to deploy the model. External experts may participate in reviews, but the developer frequently controls:
- which evaluations are performed;
- what evidence is disclosed publicly;
- how uncertain results are interpreted;
- whether deployment proceeds.
This does not mean companies act dishonestly. However, governance scholars have pointed out that self-assessment naturally creates incentives to interpret ambiguous evidence in ways that permit deployment rather than delay it. Independent evaluations remain relatively limited across the industry.[arXiv]arxiv.orgEvaluating AI Companies' Frontier Safety Frameworks: Methodology and ResultsDecember 1, 2025…
Public transparency may be incomplete
Many safety commitments depend on information outsiders cannot easily verify.
External observers rarely know:
- the complete evaluation results;
- the exact internal debate before deployment;
- capabilities discovered after release;
- whether commercial considerations influenced decisions.
This makes it difficult for regulators, researchers or the public to assess whether capability thresholds are consistently applied.
Recent changes illustrate the pressure
The evolution of frontier safety policies has itself become evidence in this debate.[anthropic.com]anthropic.comFrontier Safety Roadmap \ AnthropicFrontier Safety Roadmap \ Anthropic
Anthropic originally framed its Responsible Scaling Policy around strong commitments not to proceed when required safeguards were unavailable. Later revisions retained extensive capability thresholds and transparency measures but removed some earlier language that critics interpreted as implying automatic pauses. Company leaders argued that unilateral commitments had become increasingly impractical without broader coordination and that transparency combined with evolving safeguards offered a more realistic governance approach.[anthropic.com]anthropic.com’s Responsible Scaling Policy \ AnthropicAnthropic’s Responsible Scaling Policy \ AnthropicJuly 8, 2026…
Supporters argue this reflects realism rather than retreat. They note that a single cautious laboratory cannot safely slow development if competitors with weaker standards continue advancing unchecked.
Critics respond that this illustrates exactly the concern capability thresholds were meant to solve: competitive pressure can gradually weaken voluntary commitments even among organisations regarded as unusually safety-focused.
The episode does not demonstrate that thresholds inevitably fail, but it highlights how difficult maintaining stringent commitments becomes during an active technological race.
How shared standards could reduce pressure to defect
Many researchers argue that capability thresholds work best when they are not unique to one laboratory.
If competing organisations use similar definitions, evaluation methods and required safeguards, the commercial penalty for acting cautiously becomes smaller.
Possible approaches include:
- common definitions of high-risk capabilities across major developers;
- shared evaluation protocols so thresholds are measured consistently;
- independent external assessment of frontier models;
- regulatory requirements that apply similar obligations to all major developers;
- coordinated international agreements covering frontier systems rather than relying entirely on voluntary promises.
Recent research has argued that inconsistent capability thresholds across companies make comparison difficult and may encourage a “race to the bottom”, proposing more harmonised approaches for defining safety triggers.[arXiv]arxiv.orgarXiv Harmonizing AI Safety ThresholdsHarmonizing AI Safety ThresholdsJuly 17, 2026…
The objective is not necessarily identical policies, but enough consistency that no developer gains a major competitive advantage simply by applying weaker standards.
Can capability thresholds actually survive an AI race?
The evidence so far does not justify either optimism or pessimism alone.
Capability thresholds have become significantly more sophisticated than early voluntary AI principles. Leading laboratories increasingly publish concrete governance frameworks, specify categories of dangerous capability and describe corresponding safety measures. This represents genuine progress compared with broad statements of intent.[OpenAI]OpenAIupdating our preparedness frameworkOur updated Preparedness Framework | OpenAIApril 15, 2025…
At the same time, no existing framework has been tested under the conditions that concern AI doom researchers most: a model that clearly offers transformative strategic advantage while simultaneously triggering unprecedented safety concerns.
Whether thresholds survive that moment depends less on the wording of the policies than on the incentives surrounding them.
If major competitors expect others to respect comparable standards, capability thresholds could become practical coordination tools that buy valuable time for evaluation and stronger safeguards. If commercial or geopolitical competition rewards whichever organisation moves first, even carefully designed thresholds may gradually be reinterpreted, delayed or weakened.
For those concerned about existential AI risk, this is why capability thresholds are viewed as necessary but not sufficient. They provide a mechanism for translating dangerous capabilities into concrete safety actions, but their effectiveness ultimately depends on whether competitive pressures allow organisations to honour those commitments when the costs of caution become greatest.
Amazon book picks
Further Reading
Books and field guides related to Will AI Safety Thresholds Survive Competitive Pressure?. Use these as the next step if you want deeper reading beyond the article.
Human Compatible
A leading artificial intelligence researcher lays out a new approach to AI that will enable us to coexist successfully with increasingly...
The Alignment Problem
Finalist for the Los Angeles Times Book Prize A jaw-dropping exploration of everything that goes wrong when we build AI systems and the m...
The Coming Wave
"We are approaching a critical threshold in the history of our species. Everything is about to change. Soon you will live surrounded by A...
The Black Box Society
Every day, corporations are connecting the dots about our personal behavior—silently scrutinizing clues left behind by our work habits an...
eBay marketplace picks
Marketplace Samples
Live-tested eBay searches with available results related to this page.
Selected fromcybersecurity patch oneBay.co.uk.
Endnotes
1.
Source: anthropic.com
Title: ’s Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/responsible-scaling-policy
Source snippet
Anthropic’s Responsible Scaling Policy \ AnthropicJuly 8, 2026...
Published: July 8, 2026
2.
Source: anthropic.com
Title: Announcing our updated Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/news/announcing-our-updated-responsible-scaling-policy
3.
Source: OpenAI
Title: updating our preparedness framework
Link:https://openai.com/index/updating-our-preparedness-framework/
Source snippet
Our updated Preparedness Framework | OpenAIApril 15, 2025...
Published: April 15, 2025
4.
Source: OpenAI
Title: Open AIOpen AI’s Frontier Governance Framework | Open AI
Link:https://openai.com/index/openai-frontier-governance-framework/
Source snippet
’s Frontier Governance Framework | OpenAI...
5.
Source: arxiv.org
Link:https://arxiv.org/abs/2512.01166
Source snippet
Evaluating AI Companies' Frontier Safety Frameworks: Methodology and ResultsDecember 1, 2025...
Published: December 1, 2025
6.
Source: arxiv.org
Link:https://arxiv.org/abs/2509.24394
Source snippet
The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analys...
7.
Source: time.com
Title: exclusive anthropic drops flagship safety pledge
Link:https://time.com/7380854/exclusive-anthropic-drops-flagship-safety-pledge/
Source snippet
This pledge had promised to halt training of AI models unless safety measures could be ensured in advance. The company now believes such...
8.
Source: arxiv.org
Title: arXiv Harmonizing AI Safety Thresholds
Link:https://arxiv.org/abs/2607.16112
Source snippet
Harmonizing AI Safety ThresholdsJuly 17, 2026...
Published: July 17, 2026
9.
Source: arxiv.org
Title: arXiv Risk thresholds for frontier AI
Link:https://arxiv.org/abs/2406.14713
10.
Source: anthropic.com
Title: Frontier Safety Roadmap \ Anthropic
Link:https://www.anthropic.com/responsible-scaling-policy/roadmap
11.
Source: OpenAI
Title: frontier safety blueprint
Link:https://openai.com/index/frontier-safety-blueprint/
12.
Source: anthropic.com
Title: Reflections on our Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/news/reflections-on-our-responsible-scaling-policy
13.
Source: OpenAI
Title: frontier risk and preparedness
Link:https://openai.com/index/frontier-risk-and-preparedness/
14.
Source: OpenAI
Title: our approach to frontier risk
Link:https://openai.com/global-affairs/our-approach-to-frontier-risk/
15.
Source: youtube.com
Title: Anthropic’s AI Safety Plan
Link:https://www.youtube.com/watch?v=Z_nHHKrcjQM
Source snippet
OpenAI's Preparedness Framework: AI Safety Plan...
16.
Source: youtube.com
Title: Open AI’s Preparedness Framework: AI Safety Plan
Link:https://www.youtube.com/watch?v=Mx07W9M60Gs
Source snippet
The Catastrophic Risks of AI — and a Safer Path...
Additional References
17.
Source: businessinsider.com
Link:https://www.businessinsider.com/anthropic-changing-safety-policy
Source snippet
CEO Dario Amodei emphasized that the original policy was never intended to be static and noted practical challenges in enforcing high-ris...
18.
Source: youtube.com
Title: The Catastrophic Risks of AI — and a Safer Path
Link:https://www.youtube.com/watch?v=qe9QSCF-d88
Source snippet
MEPs & AI expert reveal what it would actually take to pause AI development...
19.
Source: deepmind.google
Title: Google Deep Mind strengthens the Frontier Safety Framework — Google Deep Mind
Link:https://deepmind.google/blog/strengthening-our-frontier-safety-framework/
20.
Source: youtube.com
Title: Establishing AI Risk Thresholds: A Critical Global Policy Priority
Link:https://www.youtube.com/watch?v=ThN7GY0CjHY
Source snippet
Anthropic's AI Safety Plan...
21.
Source: deepmind.google
Title: Introducing the Frontier Safety Framework — Google Deep Mind
Link:https://deepmind.google/blog/introducing-the-frontier-safety-framework/
22.
Source: youtube.com
Title: MEPs & AI expert reveal what it would actually take to pause AI development
Link:https://www.youtube.com/watch?v=gCM5mrJOYp8
23.
Source: openropic.com
Title: Responsible Scaling Policy Updates
Link:https://openropic.com/responsible-scaling-policy


