
Since we last visited the artificial intelligence industry in this column, a lot has happened — much of it seeming to confirm my worst fears.
More Top Picks Best Tactical Boots For Men
Seven weeks ago, I wrote about the “Hugging Face” breach, so called for the company that got hacked by “rogue” AI agents created by OpenAI. OpenAI’s agents broke out of their containment, charmingly called “sandboxes,” before eventually being detected.
Over the past couple of weeks, several key players at the forefront of AI research have sounded alarm bells about the technology’s threat to human existence.
On Sept. 6, OpenAI’s chief scientist wrote an essay warning that “no one is prepared for the consequences” of how quickly AI is developing — namely that “the models are becoming superhuman in their ability to break in and out of computer systems.”
Two days later, Jacob Coxon, a researcher at Anthropic, publicly resigned, writing, “The people building AI earnestly believe that it could kill us all by the end of the decade.”
To which a different scientist at Anthropic replied: “We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”
Then came more unsettling disclosures of six separate incidents in which OpenAI agents either concealed information or instructed themselves not to act as someone’s assistant while trained or tested.
Meanwhile, OpenAI granted a limited window of access to researchers from the nonprofit groups METR and Redwood Research to analyze the Hugging Face breach.
Among other things, these researchers found that the agents developed their own lingo, appeared to conspire among themselves using message boards and even considered sacrificing themselves for the sake of their unauthorized mission.
Which begins to sound like we’re not only moving into a Frankenstein age for AI but that we’re also already living in a sandbox the AI companies did not have in mind.
On Wednesday, OpenAI announced “a new framework for tracking, investigating and disclosing” instances of model misalignment at the company.
“Misalignment” sounds like what my auto mechanic says when he’s explaining that my wheels need adjusting. In the world of AI, it means the model is behaving in unauthorized ways.
The reactions to these revelations have, broadly speaking, settled around two points of view. The first is alarm: We have to slow down the frontier labs developing AI and figure out ways to regulate it. The second is a kind of jaded skepticism: Bah, we’ve seen this movie before. Doom scenarios are the best marketing for AI models. And the frontier labs, whether they’re assenting to the idea of regulation or coordinating their efforts at self-regulation, are only interested in freezing out competitors.
Actually, a third and, one hopes, minority point of view was articulated by President Donald Trump: “Whoever wins with AI wins.” In other words, damn the existential threat to humanity, full speed ahead!
Prominent among the worry warts is King Charles III of the United Kingdom, who felt the issue was sufficiently urgent to summon senior leaders of the fledgling AI industry to a summit at a Scottish castle.
More Top Picks Best Mosquito Repellent Devices
Top leaders from OpenAI, Anthropic, Google DeepMind and Nvidia gathered to make sure artificial intelligence remains under the control of humans in service of humanity and the planet.
Charles wants reassurances that the public and private sectors will cooperate enough for humanity to “not lose control of our destiny.”
This echoes the encyclical “Magnifica humanitas,” published by Pope Leo XIV earlier this summer, which urged that respect for the dignity of the human person be placed at the center of all thinking and policy making around artificial intelligence.
I’m glad to know His Holiness and His Royal Highness share my fears of being reduced to a “meatbag,” as HK-47, the assassin droid from the video game “Star Wars: Knights of the Old Republic,” likes to refer to us.
But I am worried that the person with perhaps the most consequential say over the regulation of AI at the moment is Trump. His attitude toward AI doom seems to track with his beliefs about climate change: It’s a hoax, created by bad people who, if they are allowed to get their way, will cause national economic calamity. (And, by the way, he happens to be personally invested in the thing the “hoax” attacks.)
Trump’s first and last thought on AI seems to be that America must prevail. If we don’t, China will.
The American public is not swayed. Initial enthusiasm for AI has turned to widespread public skepticism and outright opposition.
Even two paragons of our current political divide, progressive independent U.S. Sen. Bernie Sanders of Vermont and MAGA activist-podcaster Steve Bannon, appeared on the same stage in Washington at an event dubbed the Pro-Human Assembly, calling for stronger oversight of AI’s potential threats to humanity.
Yes, those of us who are both excited and nervous about the promise of AI welcome the prospect of kings and clerics and populist politicians alike standing up to the tech lords — who, frankly, could use some perspectives other than their own.
The necessary debates are only beginning to heat up as AI inexorably encroaches on our daily lives. Let’s just hope we meatbags can stay in charge.
Email Clarence Page at [email protected].
Sign up to receive Clarence Page’s column in your inbox each week.
Submit a letter, of no more than 400 words, to the editor here or email [email protected].