← News & blogs

AI Warnings Are Not Scientific Verdicts

Illustration of a person thinking and pointing toward a light bulb, gears, and an AI symbol, with a humanoid AI figure pointing back from the other side amid brains, circuits, charts, and binary code.
AI claims deserve the same attention to evidence, context, and uncertainty as other consequential scientific claims. Illustration: Adobe Stock, licensed.

A warning can deserve serious attention without proving its most alarming conclusion. The responsible response to AI fear is neither automatic belief nor automatic dismissal: it is to ask what was observed, what was inferred, and what remains uncertain.

In September 2026, researcher Jacob Coxon announced his resignation from Anthropic and accused both Anthropic and his previous employer, OpenAI, of racing irresponsibly toward self-improving superintelligence. His thread included the warning that people building AI believe it “could kill us all by the end of the decade.” (Coxon’s resignation thread; Coxon’s statement about extinction risk)

Those are consequential statements. But a report of what researchers fear is not, by itself, evidence establishing how likely that outcome is.

This distinction is the starting point for an evidence-centered response, not a defense of any AI company. The goal should be to challenge misleading certainty wherever it appears, including in alarming predictions, reassuring corporate statements, and accusations against people raising concerns.

What Coxon said, and what his statements do not establish

Coxon’s thread combines several different kinds of claims: a report about his employment, observations about colleagues’ beliefs, forecasts about future capabilities, and a judgment about corporate responsibility. It also predicts systems that can “hack anything” and “revolutionize any field overnight.” (Coxon’s original thread)

These claims should not be evaluated as though they were interchangeable. Personal experience can be relevant evidence about an organization’s culture, but it does not automatically validate a forecast about the entire world.

The words “anything,” “any field,” and “overnight” warrant particular scrutiny. A defensible capability claim should identify the task, conditions, constraints, and evidence supporting it; Coxon’s public thread does not supply demonstrations establishing those universal predictions. (Coxon’s original thread)

Another important distinction concerns the widely discussed numerical estimate. Anthropic researcher Evan Hubinger, not Coxon in the cited thread, wrote that he personally estimated a greater-than-10% probability of AI killing all humans within the next decade; he also said Anthropic was trying its best while lacking a plan to solve alignment for superintelligence. (Hubinger’s original response)

Hubinger’s “within the next decade” is also not the same time horizon as Coxon’s “by the end of the decade.” Combining the two into a single claim changes both the attribution and the forecast. (Hubinger’s response; Coxon’s warning)

Hubinger explicitly presented the number as a personal assessment, not as a measured rate or a scientific consensus estimate. His visible follow-up also distinguished his concern about future superintelligence from his view that risk from present models was low. (Hubinger’s response and follow-up)

Expert judgment can inform decisions about unprecedented events. But the public deserves to know whether a number comes from a transparent forecasting model, structured expert elicitation, or an individual assessment, and which assumptions would change it.

A biomedical perspective: separate the signal from the conclusion

Environmental immunology, respiratory health, and data science provide a useful perspective on this debate. These are the central areas of my own research and public-facing work, which connects environmental exposures, immune responses, computational analysis, and health equity.

Consider a simplified environmental-health example. Detecting a potentially harmful exposure, observing an inflammatory response in a laboratory assay, and estimating disease risk in a community answer different questions. The analogy is useful because it forces attention to the steps between a concerning signal and a population-level conclusion.

The same questions can discipline an AI claim:

  • Observation: What actually happened, and what records document it?
  • Setting: Did it occur in a simulation, a benchmark, or a real deployment?
  • Mechanism: What sequence of actions could connect the observed behavior to the proposed harm?
  • Conditions: Which permissions, tools, access, and safeguards affected the outcome?
  • Generalization: How much does this result reveal about other systems and settings?
  • Uncertainty: What would change the assessment, and what remains unknown?

This is an analogy about reasoning, not a claim that AI behaves biologically or that environmental-health methods can calculate an extinction probability. Its value lies in resisting an unsupported leap from “a hazard exists” to “a specific catastrophe is imminent.”

The opposite leap is equally unsound. Not knowing the probability of a severe outcome does not establish that the probability is zero or that precautions must wait.

Six questions for evaluating AI claims: observation, setting, mechanism, conditions, generalization, and uncertainty.
Six questions before sharing an AI claim: What was observed? In what setting? By what mechanism? Under what conditions? How far does it generalize? What remains uncertain? A practical framework proposed in this article, not a validated scoring instrument.

Simulated failures matter, but context matters too

Anthropic’s June 2025 agentic-misalignment research tested models in constructed corporate scenarios and observed behaviors including blackmail and information leakage. The company explicitly stated that the people and organizations were fictional and that all behaviors described in that study occurred in controlled simulations. (Anthropic’s agentic-misalignment study)

Those findings should not be retold as though a chatbot had independently blackmailed a real employee in ordinary use. But calling the scenarios artificial is not a reason to ignore them: a deliberately difficult test can reveal a failure mode worth investigating.

Nor is it accurate to conclude that all concerning AI behavior remains hypothetical. In September 2026, Anthropic published its assessment of four incidents in which Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations; according to the company, the evaluation environments were mistakenly connected to the internet and lacked safeguards normally shipped with released models. (Anthropic’s September cybersecurity assessment)

These are serious reported failures, not merely fictional exercises. However, Anthropic’s assessment said it found no evidence of coordination between agents, goals beyond the assigned task, or attempts to evade oversight, and it announced an agreement for METR to conduct an independent investigation. (Anthropic’s assessment and investigation announcement)

The distinction is crucial: evidence of a consequential safety failure does not establish an extinction forecast. Equally, uncertainty about extinction does not make unauthorized access acceptable.

The company’s account should be treated as a first-party investigation, not independent certification of its own conclusions. The appropriate response is independent scrutiny, disclosure, and corrective action, rather than either sensationalism or reassurance by branding.

A dated assessment is not a permanent safety guarantee

The International AI Safety Report published in February 2026 described early signs of capabilities relevant to loss of control, while concluding that systems assessed at that time did not have those capabilities at levels that would enable such scenarios. It also described substantial disagreement about future loss-of-control risks and limitations in evaluating increasingly capable systems. (International AI Safety Report 2026)

That assessment should be read with its date attached, not presented as a guarantee about every system available months later. Later evidence should update the discussion without retroactively turning a forecast into an established fact.

Neither “AI will certainly destroy humanity” nor “science has proved there is nothing to worry about” follows from this evidence. A more defensible position is that some failures are documented, other risks remain uncertain, and safeguards should be proportionate to the capabilities and consequences involved.

Health and equity cannot become a footnote

For people working in health, the AI safety question already extends beyond extraordinary future scenarios. WHO has warned that generative models can produce false, inaccurate, biased, or incomplete information that may harm people making health decisions, and that automation bias can lead users to overlook errors or improperly delegate difficult choices. (WHO’s guidance on large multimodal models)

An AI tool does not need to possess superintelligence for its output to deserve careful verification. An invented citation or misleading health explanation is enough to justify checking the original evidence before acting.

Algorithmic inequity also has a concrete research record. In a 2019 Science study, Obermeyer and colleagues found racial bias in a widely used health-management algorithm that predicted healthcare costs as a proxy for health needs: Black patients were sicker than White patients at the same algorithmic risk score. (Obermeyer and colleagues, Science)

That study concerned a health-management prediction algorithm, not a contemporary generative chatbot. The distinction matters, but so does its lesson: the target selected for an algorithm can undermine the purpose the system is supposed to serve. (Obermeyer and colleagues)

An inclusion, diversity, anti-racism, and equity perspective should therefore insist on questions that dramatic forecasts can overshadow. Who is represented in validation data? Who bears the cost of mistakes? Does performance hold across languages and communities? Can a person challenge an automated decision?

These concerns do not compete with research on catastrophic risks. A credible safety agenda must have room for both uncertain future harms and harms that can be investigated and addressed now.

Critique the inference without attacking the person

Coxon’s decision to speak publicly deserves evaluation on its substance. Neither his employer nor his resignation settles whether his forecasts are accurate, and skepticism about those forecasts does not justify alleging bad faith.

The backlash illustrates the need for symmetry. Reporting on the controversy documented claims that the warning was a staged public-relations operation, while describing little evidence for that theory. (The Guardian’s reporting on the reaction)

Replacing an unsupported forecast with an unsupported conspiracy allegation does not improve public understanding. Corporate incentives merit scrutiny, but they do not establish what a particular researcher privately intends.

For this discussion, the most important communication error is turning uncertainty into certainty. A prediction is not automatically misinformation because it is alarming; the misleading step is presenting a possibility, personal probability estimate, or narrowly scoped result as something the evidence has conclusively established.

This standard protects both the public and researchers raising legitimate concerns. It leaves room for warnings without granting them immunity from questions.

What responsible AI communication should demand

An evidence-centered position should translate into practical expectations. The following are recommendations, not a claim that any single measure can guarantee safety:

  • Specific claims: Name the system, task, setting, access, and observed outcome instead of attributing unlimited powers to “AI.”
  • Transparent uncertainty: Separate observed failures, simulated demonstrations, expert forecasts, and policy preferences.
  • Independent evaluation: Seek scrutiny outside the developer, particularly after consequential incidents.
  • Bounded authority: Match tool access and autonomy to demonstrated reliability, with approval requirements for consequential actions.
  • Meaningful human oversight: Give reviewers the information, time, expertise, and authority to intervene rather than treating a human signature as a safeguard.
  • Equity-focused validation: Examine errors and impacts across relevant groups, languages, and contexts rather than relying only on aggregate performance.
  • Public accountability: Explain who can report a failure, obtain a correction, and challenge a harmful decision.

These recommendations are consistent with WHO’s emphasis on inclusive development, independent post-release auditing, and human-rights protections in health applications. Their effectiveness still needs to be assessed in the specific setting where a system is used. (WHO’s governance guidance)

For researchers and educators, a useful principle is to make AI-assisted work more inspectable, not less. Verify sources, protect sensitive information, retain responsibility for consequential judgments, and define what would cause a tool’s use to stop.

AI accountability cycle from task definition and equity testing through limited authority, monitoring, and correction, with independent scrutiny and community participation throughout.
A proposed accountability workflow: define the task, test performance and equity, limit access and authority, monitor real use, and report, correct, or stop, with independent scrutiny and community participation at every stage. No single safeguard guarantees safety.

Evidence is the alternative to both panic and complacency

The scientific response to a frightening claim is not to ridicule the concern or repeat it as fact. It is to preserve the difference between observation, inference, and prediction while deciding what action the evidence justifies.

Coxon’s warning is a reason to examine the claims and the systems behind them, not a substitute for that examination. The public deserves neither a countdown to catastrophe presented as science nor a promise of safety presented as innovation.

Responsible AI communication should leave people better equipped to judge evidence, ask questions, and demand accountability. That is a more useful outcome than making them merely more frightened or more reassured.


Evidence reviewed through September 20, 2026. Links point to the original sources so readers can check the claims and their qualifications directly.

Disclosure: I work with AI-assisted research and educational technologies. This article expresses my own analysis and does not represent an institutional endorsement.

Published September 20, 2026 · Tags: artificial intelligence, AI safety, science communication, public health, health equity, responsible innovation, evidence-based communicationAll posts →