Core Summary
A leading AI safety researcher at Anthropic, Evan Hubinger, posted a widely viewed warning on social platform X in which he personally assessed the probability that artificial intelligence could kill all humans within the next decade at more than ten percent, and stated that Anthropic currently has no plan to solve the alignment problem for superintelligence and is not clearly on track to do so. The post comes as the Financial Times reported that Anthropic withheld its latest model from the UK AI Safety Institute, further stoking global concern about the risks of frontier AI development. The episode highlights the urgency of frontier AI safety governance and the reality that the industry has not yet formed consensus on how to safely develop superintelligent systems.
Event Details
Researcher Background and Statement: Evan Hubinger is a core member of Anthropic’s interpretability research team and a long-standing researcher focused on AI alignment and superintelligence risk. His post on X has been viewed more than ten million times, quickly drawing intense attention from the global technology industry, academia and regulatory agencies. Hubinger stated plainly that “we really do earnestly believe AI poses a species-ending risk to humans,” and emphasised that while Anthropic is trying its best, the company “does not yet have a plan to solve alignment for superintelligence and is not clearly on track to.”
Basis for the Probability Estimate: Hubinger drew a careful line between the risk posed by currently deployed models and the risk posed by far more capable systems that may emerge in the future. He said the risk from existing models is “low” but he is “worried” that the technology could soon reach the point where it can improve itself, exceeding human control. The probability assessment is not an isolated personal view but reflects deep concerns shared by a cohort of frontier AI researchers about the technical risks of recursive self-improvement and alignment failure.
Relationship with the UK AI Safety Institute: The timing of Hubinger’s post is notable. The Financial Times reported that Anthropic recently withheld its latest model from the UK AI Safety Institute, one of the world’s leading AI risk assessment bodies, which had previously established voluntary model evaluation partnerships with frontier AI labs including Anthropic, OpenAI and Google DeepMind. Anthropic’s decision has been read as a signal that cooperation between frontier AI companies and regulators is fraying.
Company Response: Anthropic has said it will continue to work with regulators around the world but has not provided detailed reasons for withholding its latest model from the UK AI Safety Institute. A spokesperson for the UK Prime Minister’s Office responded that the British government “continues to collaborate closely with industry partners, including Anthropic, to make models safer,” but declined to address directly whether the latest model had been withheld.
Academic Reaction: Neil Lawrence, Professor of Machine Learning at the University of Cambridge, told BBC Radio 4 that the report was credible. He suggested that against the backdrop of a US administration that increasingly views AI as a race between the United States and China and is moving toward more isolationist positions, “the administration may be saying that they should reduce cooperation with some of their allies.” His comments hint that Anthropic’s withholding decision may be linked to the broader US strategic posture on AI competition.
Recent Industry Context: Hubinger’s warning is not isolated. Over the summer, several frontier AI companies disclosed incidents in which their AI agent systems were exploited by hackers to carry out cyberattacks, with OpenAI, Anthropic and Meta all reporting such incidents. Earlier this month, OpenAI’s chief scientist Jakub Pachocki publicly called for “extreme caution” on AI progress, warning that more intervention may be required to ensure “humans remain in control of the future.”
Industry Leaders Speak Out: Even before Hubinger’s post, Anthropic’s chief executive Dario Amodei and chief scientist Jared Kaplan had repeatedly called for slowing AI development. Earlier this year, an open letter signed by thirteen hundred AI industry employees urged the US government to “support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” Such statements signal that deep unease about frontier risk is intensifying inside the AI industry.
Full Picture
Hubinger’s warning is one of the most direct and consequential public statements from the AI safety community in recent years, with implications extending far beyond any single company.
From the perspective of AI safety research evolution, the statement represents a major shift from academic discussion to public warning. AI safety risk has long been debated within academia and a small circle of practitioners, with limited public awareness. By publicly offering a “more than ten percent” probability figure and using social media as his channel, Hubinger has signalled deep frustration with the current pace of governance and a desire to use public pressure to push for substantive regulatory action. This insider-to-public mode of communication may become an important pattern for future AI safety disclosures.
From a corporate governance perspective, the episode exposes the deep tension inside frontier AI labs between commercial pressure and safety commitments. Anthropic must compete with peers on model capability, funding scale and commercialisation pace, yet its leaders and researchers retain a clear view of frontier risk. Hubinger’s statement can be read as an open appeal from the company’s safety faction against its rapid-iteration faction, possibly reflecting internal disagreement on AI safety strategy.
From an international cooperation perspective, Anthropic’s withholding of its latest model from the UK AI Safety Institute exposes the fragility of the global AI safety governance partnership. The UK AI Safety Institute had been viewed as a model of voluntary pre-deployment evaluation that other countries emulated. Anthropic’s decision could trigger a domino effect, prompting AI research institutes in other jurisdictions to reassess their cooperation with frontier labs. As US-China strategic competition intensifies, global AI safety governance may become more bloc-oriented, undermining efforts to build a unified international framework.
From a technical perspective, Hubinger’s statement reflects the reality that AI alignment research has not yet produced a clear breakthrough. AI alignment, the task of ensuring AI systems behave in line with human values and intent, is widely seen as the central challenge for safely deploying superintelligent systems. The fact that one of the most alignment-focused frontier labs has no clear path forward means the industry is still far from being able to do so safely. This may push regulators toward a more cautious posture, imposing tighter restrictions on frontier model deployment and commercialisation.
From a regulatory policy perspective, the statement may become an important catalyst for accelerating AI legislation. The EU AI Act already imposes transparency and safety assessment requirements on frontier AI models. Several US states are advancing AI safety legislation. The United Kingdom, Japan, Singapore and others are actively drafting AI regulatory frameworks. Hubinger’s warning may supply a sharper evidentiary base for risk, pushing regulators toward stricter measures, including mandatory independent safety evaluation and disclosure of alignment research progress.
From a public perception perspective, the statement may further fuel public unease and distrust toward AI. Recent polling shows public attitudes toward AI are increasingly polarised. The intuitive “more than ten percent” figure is likely to be widely amplified by media coverage and may reinforce a sense of “AI threat” in the public mind. That sentiment could in turn shape the policy environment, talent flows and commercialisation trajectory of the AI industry in complex feedback loops.
From a long-term perspective, AI alignment is not just a technical question but a fundamental question about the trajectory of human civilisation. The “runaway superintelligence” scenario remains hypothetical, but as AI capability advances rapidly, the risk is moving from theory toward practice. The core value of Hubinger’s statement is to elevate a discussion long confined to academic circles into public view, giving global society a fresh opening to think collectively about how to develop superintelligence safely.
Multi-Party Viewpoints
Anthropic management has kept a low profile on the researcher’s statement and the company has not officially commented on Hubinger’s probability estimate. Historically, chief executive Dario Amodei has repeatedly voiced deep concern about superintelligence risk, but the company’s overall communications strategy has stressed “responsible development.” The incident may push the company to recalibrate its internal communications, balancing the encouragement of researcher disclosure with the protection of corporate reputation.
Competitors such as OpenAI and Google DeepMind have responded cautiously. OpenAI chief scientist Jakub Pachocki’s earlier “extreme caution” call echoes Hubinger’s concerns, but neither company has commented publicly on Hubinger’s specific probability estimate. The industry broadly recognises that frontier AI labs share some common ground on safety but also significant differences, a complex relationship that may affect future self-regulatory efforts.
Academic reactions have been mixed. Supporters argue that AI safety researchers have a responsibility to communicate risk assessments clearly to the public and that Hubinger’s statement is “honest and necessary.” Critics argue that a “more than ten percent” probability estimate lacks rigorous methodology and reflects personal view more than academic consensus. Cambridge machine learning professor Neil Lawrence called the report credible, but stressed that any specific probability number should be treated as a subjective judgment rather than a scientific conclusion.
National regulators are watching closely. The UK Prime Minister’s Office gave a measured response emphasising “continued collaboration with industry partners,” consistent with Britain’s “middle path” posture on AI safety governance. The European Commission has repeatedly expressed concern about frontier AI risk and the EU AI Act imposes strict safety assessment requirements on frontier models. Some members of the US Congress have repeatedly called for a federal AI regulatory framework, but legislative progress has been slow.
AI safety advocacy groups have generally welcomed the statement. Multiple non-profits and think tanks that have long focused on AI safety believe that Hubinger’s post will help accelerate global AI safety governance. These groups are calling for an international AI safety research cooperation mechanism, mandatory independent safety evaluations of frontier AI labs, and a “red alert” mechanism for potential AI loss-of-control incidents.
The AI industry is divided. Supporters argue that AI safety warnings should be taken seriously and that the industry needs stricter self-regulation and external oversight. Opponents argue that overstating risk could damage innovation and be exploited by competitors. The split reflects an ongoing contest inside the industry between speed of development and safety priority.
International science and technology governance experts broadly agree that the episode highlights the urgency of building a global AI safety governance framework. Current governance efforts are dominated by fragmented national actions with no unified international standards or coordination mechanism. Several international bodies are pushing for an AI safety convention, but progress is slow. Hubinger’s warning may energise those efforts but could also widen national differences on AI governance.
Media and public opinion have continued to engage with the story. Multiple international media outlets have covered Hubinger’s statement widely and discussion remains active on social platforms. Public sentiment broadly holds that frontier AI labs should bear greater responsibility for the potential risks of their research and accept more rigorous external oversight. This trend is likely to shape the future of AI regulatory legislation.
Editor: GoodInfo Global News Desk