AI SAFETY WATCH | NEWS ANALYSIS

Michigan businesses are racing to put artificial intelligence into factories, cybersecurity systems, autonomous vehicles, health care, financial services and data centers.

But some of the researchers building the world’s most powerful AI systems are issuing an extraordinary warning: They aren’t certain they will be able to control what they’re creating.

Jacob Coxon, an AI researcher who worked at both OpenAI and Anthropic, resigned from Anthropic Tuesday and accused the companies of racing toward self-improving superintelligence while “gambling with our lives.”

Then researchers still inside Anthropic publicly backed his broader warning.

Evan Hubinger, Anthropic’s Alignment Science Lead, said he personally believes there is a greater than 10% probability AI could “kill all humans” within the next decade.

He also made an arguably more consequential admission: Anthropic doesn’t yet have a plan for ensuring that hypothetical superintelligent AI remains aligned with human objectives — and isn’t clearly on track to develop one.

Samuel Marks, another Anthropic safety researcher, said AI developers believe their technology could cause human extinction or similarly catastrophic outcomes, potentially within the next few years.

Those are extraordinary claims coming from researchers inside one of the companies developing frontier AI.

But the 10% figure requires an important qualification:

It isn’t a scientifically calculated probability.

There is no accepted scientific model capable of determining the probability that artificial intelligence will cause human extinction. Hubinger’s number is his personal assessment of an unprecedented technological risk.

The more immediate question may therefore be different:

How much authority will humans give increasingly autonomous AI systems before they know whether those systems can always be controlled?

What Exactly Are They Afraid Of?

The warning isn’t that today’s ChatGPT or Claude will suddenly become conscious, turn evil and decide to destroy humanity.

Hubinger says risks from current models remain low.

The concern is what could come next: recursive self-improvement.

Imagine an AI system capable of helping researchers build a better AI.

That system then contributes to building an even more capable successor.

If AI eventually became better than humans at AI research itself, development theoretically could accelerate faster than people could evaluate each generation.

This isn’t happening autonomously today.

But AI already is becoming deeply involved in building AI.

Anthropic reported that as of May, more than 80% of code merged into its own codebase was authored by Claude, with engineers directing and reviewing the work.

The company says its engineers now merge roughly eight times as much code per day as they did in 2024.

Anthropic explicitly cautions that recursive self-improvement isn’t inevitable. But it also warns that it “could come sooner than most institutions are prepared for.”

How Could AI Become An Existential Threat?

An AI extinction scenario generally requires several things to happen:

1. AI becomes substantially more capable. Future systems exceed today’s capabilities by enormous margins.

2. AI gains greater autonomy. Systems move from answering questions to independently executing complicated objectives.

3. AI increasingly develops AI. Artificial intelligence becomes capable of designing increasingly powerful successors.

4. Development accelerates. AI-assisted research creates new generations faster than humans can thoroughly evaluate them.

5. Humans lose reliable control. A sufficiently capable system pursues objectives in ways its creators didn’t anticipate or cannot stop.

None of those steps proves the next one will occur.

That’s why extinction probabilities remain highly uncertain.

AI Doesn’t Have To Hate Humans

One misconception surrounding AI extinction scenarios is that artificial intelligence would have to become conscious or malicious.

It wouldn’t necessarily need either.

The alignment problem is about objectives.

A sufficiently powerful autonomous system theoretically could cause enormous damage simply by pursuing a poorly specified objective extraordinarily effectively.

Humans wouldn’t have to become its enemies.

They could simply become obstacles to accomplishing its objective.

That moves the debate away from science-fiction scenarios involving machines developing hatred toward humanity.

The practical question is whether humans can reliably define what increasingly capable autonomous systems should — and should never — be allowed to do.

Why Keep Building Something You Fear?

Coxon’s answer is competition.

OpenAI, Anthropic and other AI developers are locked in an enormously expensive race to build increasingly capable models.

If one company slows down, another company — or another country — could continue.

That creates a technological dilemma:

Everyone might be safer if everyone slowed down. Nobody wants to slow down first.

Coxon argues that Anthropic understands the potential stakes but continues because it fears other developers may behave less responsibly.

The race itself could therefore become a risk multiplier.

The Nearer-Term Danger: Humans Give AI More Control

Businesses don’t have to believe AI will destroy humanity to confront the underlying problem.

AI agents increasingly can write software, operate computers, communicate with other systems and perform complicated sequences of tasks.

Every additional permission expands what an AI can accomplish — and what can happen when it makes a mistake.

The nearer-term danger therefore may not be AI suddenly seizing control.

It may be humans voluntarily giving AI increasing amounts of control because doing so improves productivity, reduces costs and provides a competitive advantage.

A company automates one process.

Then another.

Human approval disappears from routine decisions.

AI gains broader access to networks, data and operating systems.

Each decision individually could make perfectly good business sense.

The cumulative effect deserves scrutiny.

Why Michigan Businesses Should Care

Michigan sits at the center of industries where artificial intelligence increasingly interacts with the physical world.

Automakers are developing autonomous vehicles. Manufacturers are connecting AI with industrial robots. Cybersecurity companies are deploying autonomous agents. Hospitals are experimenting with AI-assisted decisions. Utilities, defense contractors and data centers operate critical infrastructure increasingly dependent on automation.

For executives, AI safety doesn’t have to mean preparing for human extinction.

It means something much more familiar:

Risk management.

Companies already restrict which employees can transfer money, access confidential information or alter critical systems.

AI agents may require similar controls: limited permissions, human approval for consequential decisions, activity logs and reliable shutdown mechanisms.

We’ve Seen This Pattern Before — With One Difference

History is filled with transformative technologies arriving before institutions were prepared for their consequences.

Industrialization preceded modern workplace-safety regulations.

Automobiles preceded comprehensive traffic laws.

Nuclear weapons preceded arms-control regimes.

Social media reached global scale before societies understood many of its consequences.

Artificial intelligence could follow that pattern.

But there’s an important difference.

Steam engines didn’t design better steam engines. Nuclear weapons didn’t help scientists design their successors.

AI increasingly can help researchers build the next generation of AI.

Whether that eventually produces uncontrollable recursive self-improvement remains unknown.

But it helps explain why researchers responsible for developing and controlling frontier AI are sounding alarms now.

The Question Nobody Can Answer

AI already is generating substantial economic benefits. It helps programmers write software, researchers analyze data, manufacturers automate processes and businesses increase productivity.

Those capabilities help explain why investment continues pouring into increasingly powerful systems.

But Coxon’s resignation exposes an uncomfortable contradiction at the center of the AI boom.

Companies developing frontier AI believe the technology could transform the global economy.

Some researchers inside those companies also believe there is a meaningful possibility it could become catastrophically dangerous.

Whether Hubinger’s extinction probability is 10%, 1%, 0.1% or effectively zero cannot currently be established scientifically.

The nearer-term question for Michigan businesses is easier.

Every time an organization gives an AI system additional authority, it makes a decision about how much human control it is willing to surrender.

The real test may not come when AI demands control.

It may come as humans willingly give it away — one efficiency gain at a time.