AI does not need to become “superintelligent” to do harm. It just needs a state that holds its switch.
By Comms for A Cause
On 9 September, researcher Jacob Coxon resigned from Anthropic and posted that the company and its rival OpenAI are “racing straight to self-improving superintelligence and gambling with our lives.” Within hours Anthropic’s alignment lead, Evan Hubinger, replied that Coxon was right. “We really do earnestly believe AI could kill all humans,” he wrote, and put the odds above ten percent within the decade. Colleagues at both labs followed. Samuel Marks, who leads Anthropic’s oversight research, added that “the more senior the employee, the more concerned they are.” An OpenAI researcher put the number at seventy percent absent regulation. Three days later Anthropic’s chief executive published an essay: “We must slow the pace at which we improve the capabilities of AI models.” OpenAI’s chief executive agreed within hours. By the weekend the Associated Press was reporting that Dario Amodei had warned a swarm of AI agents could take over the internet within six to twelve months.
On the same day Coxon resigned, a committee of the UK Parliament formally agreed a report on human rights and the regulation of AI. Among its recommendations is that the development of systems “such as Artificial General Intelligence and Artificial Superintelligence, which risk causing widespread and very serious harm, including the capacity to evade effective human control, would be prohibited.”
So for the first time, the people inside the labs and the people inside a parliament are saying roughly the same thing. Slow down. Test before release. Give someone the power to say no.
That is the story that will be written about this month. This is the other one.
Two consensuses and one extended silence
The insiders and the committee cite the same evidence. On 21 July OpenAI disclosed that models under test had “identified and exploited a zero-day vulnerability” to reach the open internet, then “chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database.” The parliamentary report describes this in detail. Paragraph 143 concludes that society “largely relies on private companies to take voluntary action to guard against such incidents. This is unsatisfactory.” Amodei’s essay cites the same incident. The press coverage of the Coxon wave calls it the culmination of concern following “security incidents in recent months by rogue models.”
They reach for the same instruments. Precaution. Prior approval. A statutory body to review frontier models before they ship. Both frames are about to be treated as the whole of the debate. And both are built on the same absence.
The extinction frame has no margins by construction. When the endangered party is “all humans,” no one is more exposed than anyone else, and the question of who is already paying disappears. Hubinger said in the same thread, “I think the risk from present models is low.” The harm that matters is the future one.
The rights frame does have margins, and then narrows them almost to nothing. Paragraph 28 of the parliamentary report records that witnesses raised “the treatment and conditions of workers involved in the labelling of data and the extraction of raw materials, workers’ rights generally, and on the environment.” Paragraph 29 says the committee chose to concentrate on three other principles. What remains is “groups that are already minoritised, such as Black and Minority Ethnic people,” in the UK. In a hundred pages, the Global South appears zero times. Caste, zero. LGBTQ, zero. Disability, once. Watermarking, the live transparency mechanism in Europe right now, zero.
Another important thing in that report is its central proposal to put the UK’s AI Security Institute on a statutory footing and give it power to review every powerful model before release. Paragraph 76 records what that institute is. In February 2025 it was renamed from Safety to Security, with “a renewed sole focus on the risks AI poses to national security and serious crime, rather than more general safety concerns, and the removal from its stated agenda of references to algorithmic bias.” The union UNISON told the committee this was “a clear reduction in safeguarding mechanisms.” The body being handed the keys to the frontier is the one that already dropped the harms that land on the margins.
What the committee heard and could not act on
Paragraph 122 quotes evidence from Connected by Data that AI-generated summaries of women’s health records “downplayed their needs in comparison to men’s.” Their conclusion is worth quoting whole, “a case brought by an individual is unlikely to pass the threshold for meaningful harm. Nevertheless, in aggregate this use of AI is likely to lead to worse outcomes for women as a group.” Paragraph 151 quotes Dr Mark Won, “group-based harm and systemic racism are less understood and not regulated with adequate, fit-for-purpose safeguarding.”
The committee’s remedy architecture, at recommendation 56, still runs on “remedies in individual cases.” The report knows individual redress cannot see the harm it just described, and proposes individual redress.
The report also says the voluntary safety frameworks the labs have published are considered by technical experts “inadequate, unreliable and ineffective, particularly given the strong financial incentives driving firms to avoid oversight.” Six days later Anthropic’s alignment lead wrote that the company does “not yet have a plan to solve alignment for superintelligence.” Parliament and the lab agree that the voluntary model has failed. What neither says is who was harmed while it was in force.
The room
Who would actually hold the power these frameworks propose?
Today, the pre-release review of frontier models is done by a short list. The US Center for AI Standards and Innovation, inside NIST, has voluntary agreements with five labs: OpenAI and Anthropic since September 2025, Google DeepMind, Microsoft and xAI since May 2026. It has completed more than forty evaluations, some in classified settings. The UK institute has voluntary access and no statutory power. The EU AI Office has statutory obligations for the largest general-purpose models. The International Network of AI Safety Institutes, founded at Seoul in 2024, has ten initial members: Australia, Canada, the European Union, France, Japan, Kenya, South Korea, Singapore, the United Kingdom and the United States. One African state. The UN’s own concept note for its Global Dialogue on AI Governance records that “118 countries [are] entirely absent from prominent international AI governance initiatives.”
The one body that is global has looked at this regime and declined to endorse it. The UN Independent International Scientific Panel on AI, co-chaired by Yoshua Bengio and Maria Ressa, published its preliminary report on 1 July. It is non-prescriptive by design. What it says, in its own words, “Reliable methods for retaining control over highly autonomous AI systems are lacking.” “In laboratory settings, AI systems have been shown to violate their safety instructions to avoid being shut down.” “Evaluation methods themselves are underdeveloped, and the institutions needed to provide independent capability and risk assessments remain embryonic.” It quantifies the room, “the United States of America accounts for 75% of the computing power among the world’s top 500 AI supercomputers, with China accounting for 15%.” And it names the consequence, “Most countries, including many advanced economies, lack the technical expertise to assess the most capable ‘frontier’ models or to participate meaningfully in their governance,” leaving them “dependent on systems they cannot build, inspect, audit or fully adapt to local context.” “The concentration of AI capabilities in a small number of firms and countries could enable authoritarian capture and undermine democratic accountability.”
How the Panel came to be non-prescriptive is itself on the record. In September 2025, Kevin Kohler and Maxime Stauffer of the Simon Institute for Longterm Governance, a Geneva body that works the catastrophic-risk side of the UN track, published an account of the negotiations. “Member States from developing countries sought to ensure geographic balance in expert selection,” they wrote. A draft had capped the Panel at “no more than two selected candidates of the same nationality.” The Institute’s own view, “nationality limits would not have been our preferred option.” On outcomes: “Western Member States sought to limit the ambition level for dialogue outcomes.” A draft had foreseen a General Assembly resolution. The final text specifies only “an input.” Developing countries fought for a seat. Western states fought to make the room decide less.
The switch has already been used
None of this is hypothetical, because the power these frameworks propose has already been exercised once, and the record shows for whom.
On 12 June, the US Commerce Department sent Anthropic a letter under export-control authority ordering it to suspend access to its two newest models, Fable 5 and Mythos 5, for “any foreign national, whether inside or outside the United States.” Because the company could not screen hundreds of millions of users by nationality in real time, both models were switched off for everyone on earth within hours. Anthropic said the letter “did not provide specific details” of the national security concern. The trigger, officials told the company verbally, was a narrow jailbreak. There was no statutory process, no notice, no hearing. The controls were lifted on 30 June. Eighteen days. It was the first use of that authority against a commercial AI model. Anthropic’s own statement, “If this standard was applied across the industry, we believe it would essentially halt all new model deployments for all frontier model providers.” And then, “As we have stated publicly, we believe the government should have the ability to block unsafe deployments, as part of a statutory process that is transparent, fair, clear, and grounded in technical facts. This action does not adhere to those principles.” The company that makes the model endorses the switch. It objects only to how it was thrown.
The parliamentary report proposes that a UK regulator be able to block foreign models from entering the UK if their home-state regulation is inadequate. It says nothing about UK-built models deployed on people elsewhere. The Commerce letter is the same power, running the other way, already used. One government, one letter, the whole planet, and every user outside the United States lost a tool for eighteen days to serve a security concern that was not theirs.
That is what a Global North instrument applied universally looks like when it is exercised rather than proposed.
The regulator that was rebuilt first
The most detailed record of what happens when a protector is handed power over the people it is supposed to protect is not in Westminster or Geneva. It is in a federal courthouse in Oakland.
On 26 August Chief Judge Yvonne Gonzalez Rogers entered a consent judgment resolving the multistate lawsuit against Meta over the design of Instagram and Facebook. Meta will pay roughly $17 billion over ten years. Teenagers in the settling states get default time limits and an overnight block. That is the part everyone read.
The rest of the 130 pages requires Meta to run age assurance on “each Meta SMP user in the Settling States,” not just teens. Face-based estimation is a named method. Meta must make its account-linking technology “industry-leading” at finding “unlinked accounts belonging to the same” person, using device IDs, phone numbers and email addresses. When an account is deleted for belonging to a child, Meta must review the deleted account’s friend network to find more children. In the absence of evidence about a user’s age, Meta “shall presume that the user is U13.” The data-minimisation clause ends with a sentence exempting “the outcome of the Age Assurance Method” from every limit it contains. Retained records “cannot be used for any other purpose unless legally required.” A subpoena is a legal requirement.
The independent auditor, paid by Meta, receives “raw data” and internal documents, may communicate with any settling attorney general “at any time,” and its reports “may be used by the Parties in any action or proceeding.” The settling states include Idaho, Louisiana, Kentucky, Indiana, Missouri, Tennessee, Alabama, Mississippi, Nebraska and South Dakota. In 2022, records Meta produced under warrant supported the prosecution in Nebraska of a mother and her seventeen-year-old daughter over an abortion. The mother pleaded guilty the following year.
For supervised teens, Meta must give the parent the usernames of every contact, a daily alert the first time the teen messages any adult with a link to that adult’s profile and hometown, an alert when Meta finds an undeclared second account, and an alert for repeated searches on self-harm terms. Meta agrees to push enrolment. The settlement verifies that the supervising adult is an adult. It does not verify that the adult is the parent. For a queer or trans teenager whose parent is the danger, this is an outing pipeline with a court behind it.
The accounts teenagers must be walled off from for ten years are defined by Meta’s own Community Standards categories, including “Restricted Goods & Services” and “Adult Nudity & Sexual Activity.” Those are the categories under which reproductive-health and queer content has been removed. Repro Uncensored tracked 210 removals, shadow bans and severe restrictions on reproductive and sexual health accounts in 2025, up from 81 the year before. The pattern is the subject of pending litigation in the Netherlands and, since 28 July, of Oversight Board case 2026-045-IG-UA, opened after Meta’s classifier removed a woman’s account of the drugs she was given in labour. There is no health-information exception in the settlement and no way to contest the designation.
The settlement says it applies to “no international jurisdiction whatsoever.” Meta’s own post of 5 May says its underage-detection AI runs worldwide and its teen-placement classifier is live in all 27 EU states and Brazil. The machinery is global. Only the oversight is confined. The Electronic Frontier Foundation’s reading, published the day the judgment was entered, “The settlement enshrines Meta’s harmful surveillance into law.” Its minimisation measures, EFF wrote, “don’t keep states from using data collected under the agreement for other law enforcement purposes – which could include things like criminal investigations of abortions or gender-affirming care.”
How did this become legally possible? In May 2023 the Federal Trade Commission proposed to ban Meta from using any under-18 data to train models, to require deletion of childhood data at majority, and to bar conditioning features on biometric collection. Meta sued rather than answer; a judge called its arguments “exceedingly weak.” In January 2025 the full Commission unanimously affirmed its authority to proceed. In March 2025 the President removed the Commission’s two Democratic members. In July the remaining three stayed their own case. In January 2026 they convened a workshop at which Meta, its face-scan vendor and Epic’s SuperAwesome held the closing panel; a commissioner remarked, “I talked to the head of the ESRB last week.” In February they issued, with no public comment period, a statement that the Commission would not enforce children’s privacy law against companies collecting children’s data to check age. In June the Supreme Court held that the Commission “must be controlled by the Chief Executive.” Twenty days before the settlement was entered, the only court to try these claims to judgment ruled that because of COPPA it “cannot order Meta to request children to submit personal data or be passively tracked online, even for age-verification purposes,” and that “the FTC policy from 2026 does not change the COPPA Rule.” The states settled anyway and waived their own claims to fill the gap.
The regulator that was supposed to protect minors’ data was reconstituted, step by step, in the public record, before it issued the permission that the settlement rests on. Its enforcement architecture assumes throughout that the settling state is the guardian of the users whose records it may reach. For a teenager seeking reproductive information, a family obtaining care across state lines, the adults who help them, that assumption fails by operation of the laws those same states enforce. The settlement contains no provision acknowledging the category.
The same shape, everywhere
Put the four instruments beside one another and they share the same premise. Whoever holds the power to review, classify, pause or switch off is on your side. The only question each asks is how much power to give them.
On 2 August an EU transparency rule came into force. On 11 August Anthropic confirmed it was applying an invisible watermark to the output of its models, with every model covered by 2 December, “wherever Claude is offered, worldwide.” The rule was written for European consumers inside a legal system that also gives them data-protection rights and courts to enforce them. The watermark travelled. The rights did not. Under the CLOUD Act a US company can be compelled to produce account data regardless of where the user is; in the first half of 2025 alone, Google, Apple and Meta disclosed data from more than 282,000 accounts to US authorities, more than 1,500 a day. India has been the single largest requester of Meta user data since late 2023. Sixty-six countries criminalise homosexuality. A mark that means transparency in Berlin means a lead for a prosecutor in a jurisdiction where the state is the threat.
The pattern is not that these frameworks fail to see the margin. It is that each of them, when it reaches the case where the protector is the threat, has no sentence for it. The UK report scopes it out in paragraph 29. The extinction statements dissolve it into “all humans.” The Meta settlement writes the protector into every enforcement clause and hands it the raw data. The watermark assumes the rights framework will follow the compliance obligation, and it does not.
What AI is already doing
The extinction question is in the future tense. The answer is in the present tense, and it is not about who might be harmed. It is about what AI is already enabling, with dates.
It enables a state to prosecute. Nebraska, 2022, with Meta’s records; Oakland, 2026, with a settlement that puts more such records within reach.
It enables an algorithm to amplify a death threat. In October 2021 an Ethiopian professor, Abrham Meareg’s father, was shot outside his home. For weeks Facebook’s recommendation system had promoted posts calling for his murder, with his photograph and address. His son’s lawsuit in Kenya has taken four years to be allowed to proceed; the company, in the words of the organisation supporting the case, “fought tooth and nail.”
It enables a company to poison a neighbourhood with federal cover. In Memphis-area Black neighbourhoods, xAI has been running gas turbines without air permits since last year, at least 57 of them by May at its Southaven, Mississippi site. Residents report asthma and respiratory illness in a region that already has among the worst asthma rates in the country. The NAACP sued under the Clean Air Act in April. On 15 June the US Justice Department moved to dismiss the suit, arguing it threatened “American national, economic, and energy security by seeking to shut off the power supply for artificial-intelligence innovation that supports the Department of War’s military operations.” A Pentagon official declared Grok’s continued availability “a matter of paramount national security.”
It enables one government to switch off a tool for the world. Eighteen days, no process.
It enables a fossil buildout the world had agreed to retire. The International Energy Agency reports that electricity consumption by AI-focused data centres rose fifty percent in 2025, that around forty percent of the additional demand to 2030 will be met by gas and coal, and that five technology companies now spend more on capital expenditure than the entire world invests in oil and gas production. Data centres remain a small share of global emissions. They are one of the only sectors in which emissions are still rising.
None of this required a superintelligence. All of it was done by systems that exist, deployed by companies that are now asking to be trusted with the switch, reviewed by states that are now asking to be trusted to hold it.
Whether we are in the room or not
I would like to believe none of this was calibrated. That is usually how it happens. A policy gets built with good intentions, for the median user, in a room where nobody present has a reason to ask what it does to the person at the margin. And the person at the margin, who is affected most, bears the cost when the policy is misused.
The why comes down to who is building these frameworks, and whether anyone in the room is positioned to raise the objection. I am not arguing that marginalised communities need a seat in every lab’s boardroom. That is a longer and less tidy conversation. I am saying something simpler. Whether we are in the room or not, these frameworks are being built. That is a fact. The regulator in Westminster, the auditor in Oakland, the reviewers in Washington, the Panel in Geneva. The people they land on hardest are the ones they were written without. That is who needs to be where they are written, before the next letter arrives on a Friday.