Within Safety Thresholds
Can AI Labs Judge Their Own Danger Fairly?
When the same lab measures risk and decides whether to release, ambiguous results may be interpreted in favour of deployment.
On this page
- How internal evaluations shape threshold decisions
- Where commercial incentives affect ambiguous results
- What independent review could verify
Page outline Jump by section
Introduction
Many frontier AI safety policies rely on a simple idea: if a model reaches a dangerous capability threshold, the developer should delay deployment or introduce stronger safeguards. The difficulty is that, in most cases today, the same laboratory that builds the model also decides whether that threshold has been crossed. This creates a potential conflict of interest. When evidence is ambiguous, commercial, competitive or reputational pressures may encourage an organisation to interpret results in ways that favour release rather than delay.
Within AI doom debates, this is not usually presented as an accusation of deliberate bad faith. Instead, it is treated as a governance problem. If capability thresholds depend primarily on internal judgement, then they may gradually shift over time without any obvious announcement. Critics describe this as “moving the goalposts”: changing the practical meaning of a safety threshold through revised interpretations, new evaluation methods or updated definitions rather than formally abandoning the policy.
How internal evaluations shape threshold decisions
Capability thresholds sound objective, but most are not simple numerical cut-offs. They depend on difficult judgements about what a model can actually do outside carefully designed tests.
For example, a laboratory may need to answer questions such as:
- Does the model genuinely automate advanced AI research, or merely assist human researchers?
- Is a successful demonstration repeatable or an isolated result?
- How much prompting or external tooling counts as part of the model’s capability?
- Should potential future improvements through fine-tuning or agent scaffolding count when deciding today’s safeguards?
These questions rarely have universally accepted answers. Different evaluation methods can produce different results, even when testing the same underlying model.
Both Anthropic’s Responsible Scaling Policy and OpenAI’s Preparedness Framework recognise that dangerous capabilities must be assessed before deployment. However, in both systems, the developer performs or commissions much of the underlying evaluation before deciding whether stronger safeguards apply. External experts may advise or review parts of the process, but they generally do not possess independent authority to declare that a threshold has been crossed.[anthropic.com]anthropic.com’s Responsible Scaling Policy \ AnthropicAnthropic’s Responsible Scaling Policy \ AnthropicJuly 8, 2026…
As a result, the most important judgement often concerns not the published threshold itself but how evidence is interpreted.
Where commercial incentives affect ambiguous results
The concern is not that laboratories necessarily falsify evaluations. Rather, ambiguous evidence can naturally be interpreted in ways that reduce disruption.
Several incentives can push in the same direction:
- delaying deployment may allow competitors to release first;
- investors and customers often expect frequent model improvements;
- public announcements create expectations around launch dates;
- internal teams may sincerely believe additional deployment experience will improve safety knowledge.
Each incentive may appear reasonable in isolation. Together, they can create systematic pressure to classify borderline cases as “not yet over the threshold”.
Behavioural research on motivated reasoning suggests that experts are not immune to unconscious bias when evaluating evidence affecting important organisational outcomes. High-stakes scientific fields have long recognised this problem by separating safety assessment from commercial decision-making where practical. Critics argue that frontier AI development increasingly faces similar institutional pressures.
This issue becomes more significant as capabilities approach poorly understood boundaries. Anthropic has publicly acknowledged that determining whether some autonomy thresholds have been crossed is becoming increasingly subjective, and has introduced additional reporting rather than relying solely on a binary internal judgement.[anthropic.com]anthropic.com’s Responsible Scaling Policy \ AnthropicAnthropic’s Responsible Scaling Policy \ AnthropicJuly 8, 2026…
How goalposts can move without changing the written policy
“Moving the goalposts” does not necessarily require rewriting a published safety framework. The practical threshold can shift through more subtle mechanisms.
Common examples include:
- Changing the evaluation benchmark. A new test may measure a different task or require stronger evidence before counting a capability as dangerous.
- Redefining success. A capability once considered sufficient to trigger safeguards may later be treated as incomplete because human supervision remains necessary.
- Adjusting prompting assumptions. Strong performance achieved through sophisticated prompting may be discounted as unrealistic, even if attackers could use similar techniques.
- Treating failures as exceptional. Successful demonstrations may be dismissed as unreliable until they occur consistently across many trials.
- Updating thresholds after capabilities improve. If standards evolve alongside rapidly improving models, the effective safety trigger may remain just beyond current systems.
Individually, each change may have legitimate technical justification. The governance concern arises when several changes consistently make deployment easier while preserving the appearance of unchanged commitments.
The problem of evaluating capabilities that are hard to measure
Some of the capabilities most relevant to existential-risk discussions are among the hardest to evaluate.
Long-term autonomous planning, strategic deception, independent scientific research or recursive improvement cannot be measured as easily as benchmark accuracy on mathematics or programming tasks.
Researchers have increasingly argued that benchmark tests alone may both underestimate and overestimate real-world capability because they simplify complex tasks into narrow scoring systems. More realistic “open-world evaluations” deliberately examine messy, long-duration activities resembling real deployment, but these evaluations require substantial human judgement and are harder to standardise.[arXiv]arxiv.orgarXiv Open-World Evaluations for Measuring Frontier AI CapabilitiesOpen-World Evaluations for Measuring Frontier AI CapabilitiesMay 19, 2026…
Ironically, the capabilities most likely to matter for AI doom scenarios may therefore be precisely those where internal interpretation plays the largest role.
Why critics argue independent review matters
For critics of industry self-governance, the central issue is verification rather than trust.
An independent reviewer could potentially:
- reproduce evaluations using different methods;
- inspect unsuccessful as well as successful test results;
- examine internal discussions about borderline cases;
- assess whether evaluation methods changed between model generations;
- verify that published thresholds were applied consistently.
This resembles safety oversight in industries such as aviation, pharmaceuticals and nuclear power, where organisations usually cannot certify their own compliance without some external scrutiny.
Supporters of stronger independent review argue that it protects companies as well as the public. If a laboratory concludes that a model remains below a threshold, outside verification can increase confidence that the decision was not influenced by commercial pressure.
Some recent governance proposals have moved modestly in this direction by incorporating external expert input or public risk reports. However, the developer typically retains ultimate authority over deployment decisions, and external reviewers often examine evidence selected by the company itself rather than conducting fully independent testing.[OpenAI]OpenAIfrontier governance framework’s Frontier Governance Framework | OpenAIMay 28, 2026…
The strongest objections to the criticism
Developers and some policy analysts raise several counterarguments.
First, frontier AI evaluations often require access to confidential model weights, training data and proprietary infrastructure. Full independent testing may therefore be technically or commercially difficult.
Second, evaluation methods evolve rapidly because model capabilities change rapidly. Updating thresholds or benchmarks is not necessarily evidence of weakening standards; it may reflect genuine scientific progress.
Third, external regulators currently possess less technical expertise and fewer resources than frontier laboratories. Internal evaluators may therefore be better placed to identify emerging risks than governments or general-purpose regulators.
Finally, excessive procedural requirements could slow safety improvements as well as commercial releases if researchers spend increasing time satisfying governance processes instead of improving evaluation methods.
These objections do not eliminate concerns about self-assessment, but they explain why replacing internal judgement with entirely external regulation has proved difficult.
What this means for AI doom arguments
Within AI doom discussions, self-assessment matters because many proposed safety frameworks assume that organisations will voluntarily stop or slow development when warning signs appear.
If the organisation deciding whether to pause also bears the costs of pausing, critics argue that safety thresholds risk becoming flexible precisely when they are under the greatest pressure. Even without deliberate misconduct, repeated judgement calls can gradually redefine what counts as “safe enough”.
Supporters of voluntary frameworks respond that today’s laboratories have incentives to protect their reputations, attract safety-conscious employees and avoid catastrophic failures. They also note that many frontier companies now publish more detailed safety documentation than was common only a few years ago.
The unresolved question is therefore not whether internal evaluations have value—they clearly do—but whether they are sufficient on their own. For readers concerned about existential risk, the credibility of capability thresholds depends not only on how carefully they are written, but on whether independent observers can verify that they continue to mean the same thing when competitive pressures intensify.
Amazon book picks
Further Reading
Books and field guides related to Can AI Labs Judge Their Own Danger Fairly?. Use these as the next step if you want deeper reading beyond the article.
The Alignment Problem
Finalist for the Los Angeles Times Book Prize A jaw-dropping exploration of everything that goes wrong when we build AI systems and the m...
The Black Box Society
Every day, corporations are connecting the dots about our personal behavior—silently scrutinizing clues left behind by our work habits an...
Weapons of Math Destruction
'A manual for the 21st-century citizen... accessible, refreshingly critical, relevant and urgent' - Financial Times 'Fascinating and deep...
Bad Blood
‘I couldn’t put down this thriller’ – Bill Gates Winner of the Financial Times/McKinsey Business Book of the Year Award. The shocking tru...
eBay marketplace picks
Marketplace Samples
Live-tested eBay searches with available results related to this page.
Selected fromartificial intelligence pin oneBay.co.uk.
Endnotes
1.
Source: anthropic.com
Title: ’s Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/responsible-scaling-policy
Source snippet
Anthropic’s Responsible Scaling Policy \ AnthropicJuly 8, 2026...
Published: July 8, 2026
2.
Source: OpenAI
Title: Open AIOur updated Preparedness Framework | Open AI
Link:https://openai.com/index/updating-our-preparedness-framework/
Source snippet
Our updated Preparedness Framework | OpenAI...
3.
Source: arxiv.org
Title: arXiv Open-World Evaluations for Measuring Frontier AI Capabilities
Link:https://arxiv.org/abs/2605.20520
Source snippet
Open-World Evaluations for Measuring Frontier AI CapabilitiesMay 19, 2026...
Published: May 19, 2026
4.
Source: OpenAI
Title: frontier governance framework
Link:https://openai.com/index/openai-frontier-governance-framework/
Source snippet
’s Frontier Governance Framework | OpenAIMay 28, 2026...
Published: May 28, 2026
5.
Source: anthropic.com
Title: Responsible Scaling Policy Version 3.0 \ Anthropic
Link:https://www.anthropic.com/news/responsible-scaling-policy-v3?e45d281a_page=1&field_format_value=3&uncat=12
6.
Source: anthropic.com
Title: ’s Transparency Hub \ Anthropic
Link:https://www.anthropic.com/transparency/voluntary-commitments
7.
Source: alignment.anthropic.com
Title: sabotage risk report
Link:https://alignment.anthropic.com/2025/sabotage-risk-report/
8.
Source: anthropic.com
Title: Responsible Scaling Policy Updates \ Anthropic
Link:https://www.anthropic.com/rsp-updates?guides=image-generation-social-good
9.
Source: OpenAI
Title: expanding on [sycophancy]({{ ‘sycophancy/’ | relative_url }})
Link:https://openai.com/index/expanding-on-sycophancy/
10.
Source: anthropic.com
Title: Sabotage evaluations for frontier models \ Anthropic
Link:https://www.anthropic.com/research/sabotage-evaluations
11.
Source: anthropic.com
Title: Announcing our updated Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/news/announcing-our-updated-responsible-scaling-policy
12.
Source: anthropic.com
Title: Reflections on our Responsible Scaling Policy \ Anthropic
Link:https://www.anthropic.com/news/reflections-on-our-responsible-scaling-policy
13.
Source: OpenAI
Title: s comment to the ntia on open model weights
Link:https://openai.com/global-affairs/openai-s-comment-to-the-ntia-on-open-model-weights/
14.
Source: OpenAI
Title: response to nist executive order on ai
Link:https://openai.com/global-affairs/response-to-nist-executive-order-on-ai/
15.
Source: OpenAI
Title: our approach to frontier risk
Link:https://openai.com/global-affairs/our-approach-to-frontier-risk/
16.
Source: OpenAI
Title: frontier risk and preparedness
Link:https://openai.com/index/frontier-risk-and-preparedness/
17.
Source: OpenAI
Link:https://openai.com/safety/how-we-think-about-safety-alignment/
18.
Source: youtube.com
Title: Anthropic’s AI Safety Plan
Link:https://www.youtube.com/watch?v=Z_nHHKrcjQM
Source snippet
OpenAI's Preparedness Framework: AI Safety Plan...
19.
Source: youtube.com
Title: Open AI’s Preparedness Framework: AI Safety Plan
Link:https://www.youtube.com/watch?v=Mx07W9M60Gs
Source snippet
New frontiers in AI governance...
20.
Source: publicnow.com
Title: Open A I Inc
Link:https://www.publicnow.com/view/E5D5932DBCEB381A6C9A571217629E8E96D98C7D
Source snippet
(via Public) / OpenAI’s Frontier Governance FrameworkMay 28, 2026 — OpenAI Inc. 05/28/2026 | News release | Distributed by Public on 05/2...
Published: May 28, 2026
21.
Source: safeguard.sh
Title: Open A I Preparedness Framework v2 Analysis
Link:https://safeguard.sh/resources/blog/openai-preparedness-framework-v2-april-2025-update
22.
Source: tracker.safer-ai.org
Link:https://tracker.safer-ai.org/company/openai/
Additional References
23.
Source: techcrunch.com
Link:https://techcrunch.com/2026/07/14/deepmind-ceo-calls-for-an-independent-standards-body-to-regulate-frontier-ai/
Source snippet
DeepMind CEO calls for an independent standards body to regulate frontier AI | TechCrunchJuly 14, 2026 — Image: Demis Hassabis, chief exe...
Published: July 14, 2026
24.
Source: aiwiki.ai
Title: Preparedness Framework (Open AI) | AI Wiki
Link:https://aiwiki.ai/wiki/preparedness_framework
Source snippet
The Preparedness Framework has been broadly received as one of the more concrete pre-deployment safety policies among frontier labs, while...
25.
Source: aievals.co
Title: Anthropic Responsible Scaling Policy
Link:https://www.aievals.co/learn/governance/anthropic-rsp
Source snippet
AI EvalsMay 29, 2026 — AI Evals Anthropic Responsible Scaling Policy 1. Learn 2. 3. Governance, Risk, Compliance 4. 5. Anthropic Responsi...
Published: May 29, 2026
26.
Source: blog.prompt20.com
Title: dangerous capability evaluations
Link:https://blog.prompt20.com/posts/dangerous-capability-evaluations/
Source snippet
prompt20.comDangerous-Capability Evaluations: How Labs Test for CBRN, Cyber, and Autonomy — Prompt20 BlogMay 31, 2026 — dangerous-capabil...
Published: May 31, 2026
27.
Source: youtube.com
Title: Anthropic Drops Hallmark Safety Pledge in Race With AI Peers
Link:https://www.youtube.com/watch?v=33lZi_Hfc8M
Source snippet
Tanya Verma - Publicly verifiable governance...
28.
Source: aisi.gov.uk
Link:https://www.aisi.gov.uk/blog/early-lessons-from-evaluating-frontier-ai-systems
29.
Source: iaps.ai
Link:https://www.iaps.ai/research/[evaluation-awareness
30.
Source: aisi.gov.uk
Link:https://www.aisi.gov.uk/blog/more-compute-more-capability-why-ai-agent-evals-need-to-account-for-test-time-compute
31.
Source: iaps.ai
Link:https://www.iaps.ai/research/responsible-scaling
32.
Source: techcrunch.com
Link:https://techcrunch.com/2026/05/07/elon-musks-lawsuit-is-putting-openais-safety-record-under-the-microscope/



