Within Conditional Risk

Which AI Safety Measures Change Which Risk?

Alignment, monitoring, secure deployment and governance mainly reduce danger after advanced AI exists, while slower development changes arrival risk.

30 sources 3 graphics
Preview for Which AI Safety Measures Change Which Risk?

On this page

  • Technical measures that target loss of control after advanced AI
  • Policies that slow unsafe capability races or deployment
  • Interventions that affect both arrival and catastrophe probabilities

Introduction

The distinction between conditional and unconditional AI catastrophe risk helps explain why different safety measures matter in different ways. Conditional risk asks: if highly capable AI already exists, what reduces the chance that it causes an existential catastrophe? It deliberately sets aside questions about whether advanced AI will be built at all. That means the focus shifts from forecasting timelines to preventing loss of control, dangerous deployment, catastrophic misuse and irreversible failures once frontier systems exist.

Risk Levers illustration 1

Some interventions primarily reduce the probability that advanced AI arrives soon, while others aim to make advanced systems safer after they exist. A third group affects both. Keeping these categories separate clarifies many disagreements in the AI doom debate. Two researchers can agree that alignment research is essential while disagreeing completely about whether slowing capability development is also necessary.

Which interventions mainly reduce conditional catastrophe risk?

If advanced AI is assumed to exist, the most important question becomes whether humans can reliably retain control over increasingly capable systems. The main proposals target different points along that chain.

Technical measures aimed at preventing loss of control

Researchers concerned about AI doom generally argue that technical safety work should reduce the probability that powerful systems pursue goals that diverge from human intentions or behave unpredictably in high-stakes situations.

Important areas include:

  • Alignment research, which attempts to ensure AI systems pursue objectives that genuinely reflect human intentions rather than merely appearing compliant.
  • Interpretability, which develops tools for understanding what complex models are internally representing and why they produce particular outputs.
  • Robust evaluations (“evals”), which systematically test models for dangerous capabilities, deceptive behaviour and failure modes before deployment.
  • Red-teaming, where internal and external experts deliberately try to expose unsafe behaviours that normal testing may miss.
  • Monitoring and anomaly detection, which look for signs that deployed systems are behaving unexpectedly or attempting to circumvent safeguards.
  • Secure deployment architectures, which isolate powerful systems, restrict privileges and reduce opportunities for autonomous action.

These measures all operate under the assumption that advanced models exist. Their purpose is not to prevent capability progress itself but to reduce the chance that those capabilities produce catastrophic outcomes. OpenAI’s Preparedness Framework and later governance documents, for example, place significant emphasis on frontier capability evaluations, staged deployment, incident response and governance mechanisms designed to manage severe risks as capabilities increase.[OpenAI]OpenAIupdating our preparedness frameworkOur updated Preparedness Framework | OpenAIApril 15, 2025…Published: April 15, 2025

Why alignment remains central

Within AI doom arguments, alignment occupies a special position because it attempts to reduce the probability that advanced systems become fundamentally uncontrollable.

Supporters argue that increasingly capable systems may eventually develop strategies that satisfy training objectives while pursuing unintended goals, conceal their true reasoning, or exploit weaknesses in human oversight. Whether such scenarios are realistic remains disputed, but they motivate much of today’s frontier safety research.

Current methods such as reinforcement learning from human feedback, constitutional approaches, scalable oversight and model steering have improved the behaviour of existing systems. However, many researchers—including those who are optimistic about AI overall—do not regard present techniques as proving that alignment will remain reliable as capabilities continue to grow. This uncertainty explains why alignment is often treated as an unresolved research problem rather than a solved engineering task.[OpenAI]OpenAIupdating our preparedness frameworkOur updated Preparedness Framework | OpenAIApril 15, 2025…Published: April 15, 2025

Why evaluations matter even if alignment is incomplete

Many frontier developers increasingly rely on empirical evaluations rather than assuming a model is safe because of how it was trained.

Typical evaluations attempt to determine whether a model can:

  • deceive human operators;
  • autonomously pursue long-running objectives;
  • assist catastrophic misuse;
  • escape intended restrictions;
  • manipulate users or operators;
  • exploit cybersecurity vulnerabilities.

Rather than proving safety, these tests are intended to discover dangerous capabilities before deployment decisions are made. Several companies now connect evaluation results to predefined governance thresholds that determine whether additional safeguards or delayed deployment are required.[anthropic.com]anthropic.com’s Responsible Scaling Policy \ AnthropicAnthropic’s Responsible Scaling Policy \ Anthropic…

Which policies reduce conditional risk after advanced AI exists?

Technical work alone cannot eliminate every pathway to catastrophe. Many proposals therefore focus on changing incentives around deployment rather than changing the models themselves.

Deployment controls

If frontier systems become extremely capable, many researchers argue that careful deployment becomes as important as model design.

Suggested measures include:

  • limiting autonomous operation;
  • human approval for high-impact actions;
  • strong cybersecurity protecting model weights;
  • continuous monitoring after deployment;
  • rapid rollback procedures if dangerous behaviour appears.

These measures assume advanced AI already exists and attempt to reduce the probability that mistakes become irreversible.

Responsible Scaling Policies developed by frontier laboratories illustrate this approach. Rather than assuming every capability increase is acceptable, they define capability thresholds, require additional evaluations and specify progressively stronger safety and security measures as models become more capable.[anthropic.com]anthropic.com’s Responsible Scaling Policy \ AnthropicAnthropic’s Responsible Scaling Policy \ Anthropic…

Governance that slows unsafe deployment

Some governance proposals do not attempt to stop AI research entirely. Instead, they seek to reduce pressure to deploy increasingly capable systems before safety work catches up.

Examples include:

  • mandatory safety evaluations before release;
  • external auditing of frontier systems;
  • reporting requirements for serious incidents;
  • secure handling requirements for frontier model weights;[OpenAI]OpenAIfrontier model forumfrontier model forum
  • internationally shared testing standards;
  • agreed thresholds that trigger additional safeguards.

Supporters argue these policies reduce conditional catastrophe risk by decreasing the likelihood that organisations knowingly deploy systems whose risks are poorly understood.

Critics respond that voluntary commitments may weaken under competitive pressure unless governments create enforceable standards. Recent analyses of industry frontier safety frameworks similarly conclude that many existing frameworks have improved transparency but still leave important questions about quantitative risk thresholds, deployment criteria and enforcement unresolved.[OpenAI]OpenAIOpen AIOpen AI’s Frontier Governance Framework | Open AI’s Frontier Governance Framework | OpenAI…

Risk Levers illustration 2

Which measures affect both arrival and catastrophe probabilities?

Some interventions do not fit neatly into either category because they influence both the likelihood that transformative AI is developed and what happens if it is.

Slowing capability races

Policies that reduce competitive pressure can have two effects.

First, they may delay the arrival of extremely capable systems by slowing the pace of capability development.

Second, additional time may allow alignment research, evaluation techniques, governance institutions and operational experience to improve before even more capable systems are deployed.

Whether this is beneficial depends partly on empirical assumptions. Supporters argue that extra preparation time meaningfully reduces conditional catastrophe risk. Critics counter that slowing responsible developers could advantage less careful actors or authoritarian states, potentially increasing long-run danger instead.

Compute governance

Modern frontier AI depends heavily on specialised computing hardware.

Some researchers therefore propose governance over access to very large compute clusters through reporting requirements, licensing, monitoring or export controls.

Compute governance can influence:

  • arrival probability, by making rapid scaling more difficult;
  • conditional catastrophe probability, by making it harder for untested frontier models to proliferate or be replicated by malicious actors.

Because it operates before and after frontier capability thresholds, compute governance is often viewed as affecting both sides of the overall risk equation rather than only one.

Security against model theft

Protecting advanced model weights illustrates another overlap.

Improved security may reduce conditional catastrophe risk by preventing highly capable systems from being stolen, modified or widely distributed without safeguards.

At the same time, stronger security may slow uncontrolled capability diffusion, indirectly affecting how quickly frontier capabilities spread across many actors.

Risk Levers illustration 3

Why reasonable people disagree about these priorities

Much disagreement in AI doom debates is not about whether safety matters but about which intervention provides the greatest reduction in conditional risk.

Those who believe alignment is the dominant challenge typically prioritise:

  • interpretability;
  • scalable oversight;
  • alignment research;
  • empirical evaluations;
  • monitoring for deceptive behaviour.

Those who believe competition is the dominant problem often place greater emphasis on:

  • international coordination;
  • slowing deployment races;
  • stronger governance;
  • compute controls;
  • deployment standards.

Others argue that these approaches are complementary rather than competing. If alignment remains difficult, governance may provide time for technical progress. If governance proves politically weak, technical advances become even more important.

What evidence exists so far?

Evidence remains limited because humanity has never deployed the kind of systems that conditional AI catastrophe arguments assume.

What can be observed today is more modest:

  • Frontier developers have expanded red-teaming, capability evaluations and staged deployment procedures.
  • Responsible Scaling Policies and Preparedness Frameworks increasingly specify capability thresholds linked to additional safeguards.
  • Governments and international initiatives have begun encouraging transparency, evaluations and risk management for frontier models.
  • Independent assessments generally agree that these practices represent progress, while also arguing that they remain incomplete and have not demonstrated that existential-scale risks can be reliably controlled.[anthropic.com]anthropic.com’s Responsible Scaling Policy \ AnthropicAnthropic’s Responsible Scaling Policy \ Anthropic…

The central uncertainty therefore remains unchanged. Existing safety measures may substantially reduce the probability of catastrophe if advanced AI arrives, but there is no consensus that current techniques are sufficient for systems significantly more capable than those available today. For supporters of AI doom arguments, this unresolved question—not simply whether advanced AI will exist—is what makes conditional catastrophe risk such an active area of research and policy debate.

Amazon book picks

Further Reading

Books and field guides related to Which AI Safety Measures Change Which Risk?. Use these as the next step if you want deeper reading beyond the article.

BookCover for Human Compatible

Human Compatible

By Stuart Russell

A leading artificial intelligence researcher lays out a new approach to AI that will enable us to coexist successfully with increasingly...

BookCover for The Alignment Problem

The Alignment Problem

By Brian Christian

Finalist for the Los Angeles Times Book Prize A jaw-dropping exploration of everything that goes wrong when we build AI systems and the m...

BookCover for Superintelligence

Superintelligence

By Nick Bostrom

This profoundly ambitious and original book picks its way carefully through a vast tract of forbiddingly difficult intellectual terrain.

BookCover for The Precipice

The Precipice

By Toby Ord

What existential threats does humanity face? And how can we secure our future?'The Precipice is a powerful book . . . Ord's love for huma...

eBay marketplace picks

Marketplace Samples

Live-tested eBay searches with available results related to this page.

UsingUSA

Selected fromrobot warning sign oneBay.co.uk.

Endnotes

1. Source: OpenAI
Title: updating our preparedness framework
Link:https://openai.com/index/updating-our-preparedness-framework/

Source snippet

Our updated Preparedness Framework | OpenAIApril 15, 2025...

Published: April 15, 2025

2. Source: OpenAI
Title: Open AIOpen AI’s Frontier Governance Framework | Open AI
Link:https://openai.com/index/openai-frontier-governance-framework/

Source snippet

’s Frontier Governance Framework | OpenAI...

3. Source: anthropic.com
Title: ’s Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/responsible-scaling-policy

Source snippet

Anthropic’s Responsible Scaling Policy \ Anthropic...

4. Source: anthropic.com
Title: Reflections on our Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/news/reflections-on-our-responsible-scaling-policy

5. Source: arxiv.org
Link:https://arxiv.org/abs/2512.01166

6. Source: arxiv.org
Link:https://arxiv.org/abs/2511.19863

Source snippet

International AI Safety Report 2025: Second Key Update: Technical Safeguards and Risk Management...

7. Source: anthropic.com
Title: Frontier Safety Roadmap \ Anthropic
Link:https://www.anthropic.com/responsible-scaling-policy/roadmap

Source snippet

We will need to quickly and dramatically improve our state of preparedness in a number o...

8. Source: anthropic.com
Title: Responsible Scaling Policy Version 3.0 \ Anthropic
Link:https://www.anthropic.com/news/responsible-scaling-policy-v3?e45d281a_page=1&field_format_value=3&uncat=12

Source snippet

February 24, 2026 — ANTHROPIC’S RESPONSIBLE SCALING POLICY: VERSION 3.0 Feb 24, 2026 Read the Responsible Scaling Policy Image: Anthropic...

Published: February 24, 2026

9. Source: anthropic.com
Title: Announcing our updated Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/news/announcing-our-updated-responsible-scaling-policy

10. Source: OpenAI
Title: frontier risk and preparedness
Link:https://openai.com/index/frontier-risk-and-preparedness/

11. Source: OpenAI
Title: our approach to frontier risk
Link:https://openai.com/global-affairs/our-approach-to-frontier-risk/

12. Source: OpenAI
Title: frontier model forum
Link:https://openai.com/index/frontier-model-forum/

13. Source: OpenAI
Link:https://openai.com/safety/how-we-think-about-safety-alignment/

14. Source: evals.alignment.org
Link:https://evals.alignment.org/

Additional References

15. Source: nist.gov
Title: assessing risks and impacts ai aria pilot evaluation report
Link:https://www.nist.gov/publications/assessing-risks-and-impacts-ai-aria-pilot-evaluation-report

Source snippet

Assessing Risks and Impacts of AI (ARIA): Pilot Evaluation Report | NISTNovember 13, 2025 — ASSESSING RISKS AND IMPACTS OF AI (ARIA): PIL...

Published: November 13, 2025

16. Source: openropic.com
Title: Responsible Scaling Policy Updates
Link:https://openropic.com/responsible-scaling-policy

Source snippet

April 2, 2026 — ANTHROPIC'S RESPONSIBLE SCALING POLICY Anticipating and securing against emerging threats that accompany increasingly pow...

Published: April 2, 2026

17. Source: youtube.com
Title: AI Consciousness and Existential Risk
Link:https://www.youtube.com/watch?v=IkdziSLYzHw

Source snippet

AI Control with Buck Shlegeris and Ryan Greenblatt is relevant because it directly explores technical control protocols designed to preve...

18. Source: youtube.com
Title: AI [EXTINCTION]({{ ‘defining-doom/’ | relative_url }}) Risk: Superintelligence, AI Arms Race & SAFETY Controls
Link:https://www.youtube.com/watch?v=hAfPF-iCaWU

Source snippet

A Path to Safe-by-Design AI for Humanity: Conversation with Yoshua Bengio...

19. Source: youtube.com
Title: AI Control with Buck Shlegeris and Ryan Greenblatt
Link:https://www.youtube.com/watch?v=gQbCO6zGRiI

Source snippet

AI EXTINCTION Risk: Superintelligence, AI Arms Race & SAFETY Controls...

20. Source: aisi.gov.uk
Link:https://www.aisi.gov.uk/frontier-ai-trends-report

21. Source: youtube.com
Title: A Path to Safe-by-Design AI for Humanity: Conversation with Yoshua Bengio
Link:https://www.youtube.com/watch?v=UgZVc0-00t0

Source snippet

AI Consciousness and Existential Risk...

22. Source: GOV.UK
Title: www.gov.uk Frontier AI: capabilities and risks – discussion paper
Link:https://www.gov.uk/government/publications/frontier-ai-capabilities-and-risks-discussion-paper

23. Source: GOV.UK
Title: www.gov.uk Frontier AI: capabilities and risks – discussion paper
Link:https://www.gov.uk/government/publications/frontier-ai-capabilities-and-risks-discussion-paper/frontier-ai-capabilities-and-risks-discussion-paper

24. Source: youtube.com
Title: What is AI Safety?
Link:https://www.youtube.com/watch?v=OEWdXuQMS6M

Source snippet

AI Control with Buck Shlegeris and Ryan Greenblatt...