Within Safety Thresholds

Did Competitive Pressure Weaken Anthropic's Safety Commitments?

Anthropic's changing safety commitments offer a concrete test of whether voluntary capability thresholds can survive intense competition.

30 sources 3 graphics
Preview for Did Competitive Pressure Weaken Anthropic's Safety Commitments?

On this page

  • What changed in the Responsible Scaling Policy
  • Why supporters call the revisions pragmatic
  • Why critics see a warning for voluntary safeguards

Introduction

Anthropic has often been presented as the frontier AI company most willing to make explicit safety commitments about potentially dangerous systems. That makes revisions to its Responsible Scaling Policy (RSP) especially important in debates about AI doom and existential risk. If the company most closely associated with voluntary safety guardrails concludes that some commitments are impractical under competitive conditions, critics argue this may reveal a structural weakness in voluntary governance rather than an isolated policy change.

Anthropic Case illustration 1

The significance of the case is not simply whether Anthropic became more or less cautious. Rather, it is whether intense commercial and strategic competition changes what even safety-focused organisations believe they can realistically promise. Within the broader question of whether capability thresholds can survive an AI race, Anthropic provides one of the clearest real-world tests.

What changed in the Responsible Scaling Policy?

Anthropic introduced the first version of its Responsible Scaling Policy in September 2023 as a framework linking increasingly dangerous AI capabilities to progressively stronger security and governance measures. Rather than relying on vague commitments, the policy established AI Safety Levels (ASLs), capability thresholds and associated safeguards for catastrophic risks such as advanced cyber capabilities, biological misuse and autonomous AI research.[anthropic.com]anthropic.comAnnouncing Anthropic's Responsible Scaling Policy \ AnthropicSeptember 19, 2023…Published: September 19, 2023

A major revision arrived in October 2024. Anthropic expanded and refined its capability thresholds, introduced a more formal “safety case” style of reasoning borrowed from high-risk industries, and clarified that stronger safeguards would be triggered when specified dangerous capabilities were reached rather than according to overall model size or general intelligence. The company presented these changes as making the policy more operational rather than weakening it.[anthropic.com]anthropic.comAnnouncing our updated Responsible Scaling Policy \ AnthropicOctober 15, 2024…Published: October 15, 2024

The most controversial revisions came with Version 3.0 in February 2026.[anthropic.com]anthropic.comResponsible Scaling Policy Version 3.0 \ AnthropicFebruary 24, 2026…Published: February 24, 2026

Earlier versions had been widely interpreted as containing an unusually strong commitment that Anthropic would not continue developing or deploying models unless appropriate safeguards were already in place. Version 3.0 instead shifted towards an ongoing risk-management model built around published Frontier Safety Roadmaps, regular Risk Reports and iterative governance processes. Rather than emphasising categorical pauses, the revised policy focused on documenting identified risks, required mitigations and continuing evaluation as capabilities advanced. Anthropic described the rewrite as reflecting lessons learned from more than two years of operating the framework while improving transparency and accountability.[anthropic.com]anthropic.comResponsible Scaling Policy Version 3.0 \ AnthropicFebruary 24, 2026…Published: February 24, 2026

This distinction matters because critics argue that replacing an explicit stopping commitment with more discretionary governance changes the practical force of capability thresholds, even if many technical evaluation requirements remain.

13:52

Why supporters call the revisions pragmatic

Supporters of the revisions argue that frontier AI development had changed enough to expose weaknesses in the original policy.

One argument is that dangerous capabilities are proving harder to identify through simple threshold tests than originally expected. Anthropic itself has acknowledged that determining whether models cross some autonomy-related thresholds increasingly involves subjective judgement rather than clear-cut measurements. The company has therefore expanded documentation requirements, including public Risk Reports, instead of relying entirely on binary threshold decisions.[anthropic.com]anthropic.com’s Responsible Scaling Policy \ AnthropicAnthropic’s Responsible Scaling Policy \ Anthropic…

A second argument concerns competitive dynamics.

Anthropic executives have argued that a unilateral commitment to halt development would not necessarily reduce global catastrophic risk if competitors continued developing comparable systems. According to this reasoning, voluntarily removing one safety-conscious laboratory from the frontier could simply leave less cautious organisations setting the pace. The company therefore argues that combining continued development with extensive evaluations, transparency reports and stronger governance mechanisms may produce better real-world outcomes than rigid promises that only one organisation follows.[anthropic.com]anthropic.comResponsible Scaling Policy Version 3.0 \ AnthropicFebruary 24, 2026…Published: February 24, 2026

Supporters also note that the revised framework introduced several governance features absent from the original policy, including more structured public reporting, external review mechanisms, detailed Frontier Safety Roadmaps and ongoing revisions as scientific understanding improves. From this perspective, the policy became more operational even if some headline commitments became less absolute.[anthropic.com]anthropic.comResponsible Scaling Policy Version 3.0 \ AnthropicFebruary 24, 2026…Published: February 24, 2026

Anthropic Case illustration 2

Why critics see a warning for voluntary safeguards

For many researchers concerned about AI doom, however, the symbolic importance of the revisions outweighs the technical details.

Anthropic had frequently been cited as evidence that frontier laboratories could voluntarily commit themselves to slowing or stopping development if catastrophic-risk thresholds were crossed. Weakening or removing language interpreted as requiring such pauses therefore appears, to critics, as evidence that commercial incentives eventually reshape even the strongest voluntary commitments.[techradar.com]techradar.comCritics argue the shift demonstrates the limitations of voluntary industry commitments. Despite advocating for regulation—evident in Anth…

Critics also point to a more general governance concern.

Capability thresholds only matter if they create decisions that companies would otherwise prefer not to make. If thresholds are repeatedly revised, interpreted more flexibly or accompanied by increasing managerial discretion, then the practical constraint may weaken precisely when financial and strategic pressure becomes greatest.

This does not necessarily imply bad faith. Many observers instead frame the issue as a coordination problem. A laboratory may genuinely believe delaying deployment is safer while simultaneously believing that delaying alone merely hands leadership to rivals without reducing overall global risk. Under those incentives, voluntary commitments naturally evolve towards approaches that permit continued competition.

This is one reason many governance researchers argue that industry self-regulation should eventually be supplemented by external oversight or shared international standards. If every organisation faces the same mandatory requirements, individual firms face less pressure to choose between safety commitments and competitive position.[arXiv]arxiv.orgarXiv Intolerable Risk Threshold Recommendations for Artificial IntelligenceIntolerable Risk Threshold Recommendations for Artificial IntelligenceMarch 4, 2025…Published: March 4, 2025

22:39

What the case reveals about competitive pressure

The Anthropic case illustrates several broader lessons about capability thresholds during an AI race.

First, competitive pressure does not necessarily eliminate safety policies. Anthropic continues to publish one of the most detailed public governance frameworks among frontier AI developers, and it has repeatedly revised rather than abandoned the Responsible Scaling Policy.[anthropic.com]anthropic.com’s Responsible Scaling Policy \ AnthropicAnthropic’s Responsible Scaling Policy \ Anthropic…

Second, competition appears to influence the type of commitments organisations consider sustainable. Broad promises to stop development have increasingly given way to commitments centred on transparency, evaluation, documentation and iterative risk management.

Third, the case demonstrates that capability thresholds are not purely technical. Whether a model has crossed a threshold depends on evaluation methods, interpretation of uncertain evidence and organisational judgement. Those judgements are inevitably made within commercial and geopolitical contexts rather than in isolation.

Finally, the episode strengthens one of the central arguments made by many AI doom researchers: the effectiveness of capability thresholds cannot be judged solely by reading policy documents. The crucial question is whether organisations maintain costly commitments when market leadership, investor expectations and strategic competition create strong incentives to reinterpret them.

Anthropic Case illustration 3

What this means for debates about AI doom

Within existential-risk discussions, Anthropic’s policy revisions are interpreted in sharply different ways because they support different theories about how AI governance works.

Those more optimistic about voluntary governance argue that the revisions show a company learning from practical experience while preserving substantial investment in safety evaluations, public reporting and catastrophic-risk analysis. On this reading, adaptable governance is more credible than rigid promises that may later prove unrealistic.

Those more pessimistic about AI doom prevention see the opposite lesson. They argue that if the laboratory most closely associated with voluntary safety commitments concluded that categorical restraints were difficult to sustain under competitive pressure, then reliance on voluntary capability thresholds alone is unlikely to remain reliable as models become more economically and strategically valuable.

The Anthropic case therefore does not settle whether capability thresholds can survive an AI race. Instead, it provides one of the strongest available examples of the central dilemma: safety commitments are easiest to make before they become genuinely expensive. The real test comes when honouring them risks falling behind competitors.

10:07

Amazon book picks

Further Reading

Books and field guides related to Did Competitive Pressure Weaken Anthropic's Safety Commitments?. Use these as the next step if you want deeper reading beyond the article.

BookCover for The Coming Wave

The Coming Wave

By Mustafa Suleyman

"We are approaching a critical threshold in the history of our species. Everything is about to change. Soon you will live surrounded by A...

BookCover for Human Compatible

Human Compatible

By Stuart Russell

A leading artificial intelligence researcher lays out a new approach to AI that will enable us to coexist successfully with increasingly...

BookCover for The Alignment Problem

The Alignment Problem

By Brian Christian

Finalist for the Los Angeles Times Book Prize A jaw-dropping exploration of everything that goes wrong when we build AI systems and the m...

BookCover for Power and Progress

Power and Progress

By Daron Acemoglu, Simon Johnson

A bold reinterpretation of economics and history revealing why technology does not inevitably lead to shared prosperity, and how we must...

eBay marketplace picks

Marketplace Samples

Live-tested eBay searches with available results related to this page.

UsingUSA

Selected fromAnthropic sticker oneBay.co.uk.

Endnotes

1. Source: anthropic.com
Title: Announcing Anthropic’s Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/news/anthropics-responsible-scaling-policy

Source snippet

September 19, 2023...

Published: September 19, 2023

2. Source: anthropic.com
Title: Announcing our updated Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/news/announcing-our-updated-responsible-scaling-policy

Source snippet

October 15, 2024...

Published: October 15, 2024

3. Source: anthropic.com
Title: ’s Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/responsible-scaling-policy

Source snippet

Anthropic’s Responsible Scaling Policy \ Anthropic...

4. Source: anthropic.com
Title: Responsible Scaling Policy Version 3.0 \ Anthropic
Link:https://www.anthropic.com/news/responsible-scaling-policy-v3?e45d281a_page=1&field_format_value=3&uncat=12

Source snippet

February 24, 2026...

Published: February 24, 2026

5. Source: techradar.com
Link:https://www.techradar.com/ai-platforms-assistants/anthropic-drops-its-signature-safety-promise-and-rewrites-ai-guardrails

Source snippet

Critics argue the shift demonstrates the limitations of voluntary industry commitments. Despite advocating for regulation—evident in Anth...

6. Source: arxiv.org
Title: arXiv Intolerable Risk Threshold Recommendations for Artificial Intelligence
Link:https://arxiv.org/abs/2503.05812

Source snippet

Intolerable Risk Threshold Recommendations for Artificial IntelligenceMarch 4, 2025...

Published: March 4, 2025

7. Source: arxiv.org
Title: arXiv Taking control: Policies to address [extinction]({{ ‘defining-doom/’ | relative_url }}) risks from advanced AI
Link:https://arxiv.org/abs/2310.20563

8. Source: governance.ai
Title: anthropics rsp v3 0 how it works whats changed and some reflections
Link:https://www.governance.ai/analysis/anthropics-rsp-v3-0-how-it-works-whats-changed-and-some-reflections

9. Source: anthropic.com
Title: ’s Transparency Hub \ Anthropic
Link:https://www.anthropic.com/transparency?e45d281a_page=5

10. Source: anthropic.com
Title: ’s Transparency Hub \ Anthropic
Link:https://www.anthropic.com/transparency/voluntary-commitments/security%26privacy

11. Source: anthropic.com
Title: ’s Transparency Hub \ Anthropic
Link:https://www.anthropic.com/transparency/voluntary-commitments

12. Source: alignment.anthropic.com
Title: sabotage risk report
Link:https://alignment.anthropic.com/2025/sabotage-risk-report/

13. Source: anthropic.com
Title: Activating AI Safety Level 3 protections \ Anthropic
Link:https://www.anthropic.com/news/activating-asl3-protections?subjects=alignment

14. Source: anthropic.com
Title: Responsible Scaling Policy Updates \ Anthropic
Link:https://www.anthropic.com/rsp-updates?guides=image-generation-social-good

15. Source: anthropic.com
Title: The case for targeted regulation \ Anthropic
Link:https://www.anthropic.com/news/the-case-for-targeted-regulation

16. Source: anthropic.com
Title: Reflections on our Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/news/reflections-on-our-responsible-scaling-policy

17. Source: tracker.safer-ai.org
Link:https://tracker.safer-ai.org/company/anthropic/

Additional References

18. Source: pcgamer.com
Link:https://www.pcgamer.com/software/ai/anthropic-ditches-its-defining-safety-promise-to-pause-dangerous-ai-development-because-its-basically-pointless-when-everybody-else-is-blazing-ahead/

Source snippet

Previously, under its Responsible Scaling Policy (RSP), Anthropic pledged to halt AI development should new systems reach dangerous capab...

19. Source: youtube.com
Title: Anthropic CEO warns that without guardrails, AI could be on dangerous path
Link:https://www.youtube.com/watch?v=aAPpQC-3EyE

Source snippet

Anthropic Responsible Scaling Policy safety Zac Hatfield-Dodds | Anthropic’s Responsible Scaling Policy @ Vision Weekend US 2024 Foresigh...

20. Source: openropic.com
Title: Responsible Scaling Policy Updates
Link:https://openropic.com/responsible-scaling-policy

Source snippet

April 2, 2026 — ANTHROPIC'S RESPONSIBLE SCALING POLICY Anticipating and securing against emerging threats that accompany increasingly pow...

Published: April 2, 2026

21. Source: youtube.com
Link:https://www.youtube.com/watch?v=9IhcygeoKRs

Source snippet

Anthropic Vs. OpenAI: How Safety Became The Advantage In AI...

22. Source: youtube.com
Title: Anthropic Vs. Open AI: How Safety Became The Advantage In AI
Link:https://www.youtube.com/watch?v=JILSzhssMsk

Source snippet

Anthropic CEO warns that without guardrails, AI could be on dangerous path...

23. Source: youtube.com
Title: Anthropic Drops Hallmark Safety Pledge in Race With AI Peers
Link:https://www.youtube.com/watch?v=33lZi_Hfc8M

Source snippet

Anthropic Loosens Safety Pledge as AI Race Tightens...

24. Source: youtube.com
Title: Anthropic’s AI safety contradiction
Link:https://www.youtube.com/watch?v=weZ5HQiwabE

Source snippet

Zac Hatfield-Dodds | Anthropic’s Responsible Scaling Policy @ Vision Weekend US 2024...

25. Source: safer-ai.org
Title: anthropics responsible scaling policy update makes a step backwards
Link:https://www.safer-ai.org/anthropics-responsible-scaling-policy-update-makes-a-step-backwards

26. Source: policywindow.org
Title: Anthropic Responsible Scaling Policy (RSP) v2 | Policy Window
Link:https://www.policywindow.org/wiki/anthropic-rsp

27. Source: youtube.com
Title: Anthropic Loosens Safety Pledge as AI Race Tightens
Link:https://www.youtube.com/watch?v=DJdK-ndXPiE

Source snippet

Anthropic's AI safety contradiction...