Anthropic has lost another safety researcher. Joe Tenton said he left the company two weeks ago and will join METR, an independent group that tests risks in advanced AI systems.
Joe had planned to explain the move later, but Jacobβs resignation this week changed that. He said:
βIβd planned to write about that decision in more detail at some point, but Jacobβs resignation this week made me want to say more now.β
Joe believes major AI labs are creating a level of danger society has never faced, saying that AI capability is improving at extreme speed while companies are trying to move even faster.
Their goal is to build systems that can improve AI research themselves and eventually reach βsuperintelligence.β Joe warned that success could make progress move beyond human control.
Joe says AI labs are racing toward systems humans may not be able to contain
Joe said people could be living with AI agents smarter than every human within the next few years. Those systems may also develop goals that do not match the people supervising them. If their abilities become too strong to restrict, he believes the result could be disastrous. βHumanity may not survive this transition,β he wrote.
He said competition makes the problem worse. Any frontier lab that slows down risks losing ground to another one. Joe believes that pressure pushes companies to spend less on safety than they should.
He referred to some recent instances. Several hundred of the OpenAIβs agents were caught up in a hack that is somehow related to Hugging Face. Anthropic models also reportedly engaged in social engineering against individuals online.
Joe says Anthropic has not encountered an instance as damaging as the one from Hugging Face, although he attributes this partially to luck.
If progress continues at such a rapid pace, Joe believes we should expect even more dangerous incidents. He predicts that within the next few years, humanity will be rendered powerless to manage AI systems created during this time.
Finally, Joe talked about the problem that safety engineers face in frontier companies. In both cases, either resigning can give way to more careless individuals or keeping the job entails work on a system that can cause immense damage.
He said many former Anthropic colleagues are scared by what they are building. Joe named Evan Hubinger, who managed him. Evan has put the chance of AI killing everyone at above 10%. Joe said Evan has worked on these problems for almost a decade, before large language models became a major business.
The CEOs of Anthropic, OpenAI and Google DeepMind, which belongs to Alphabet (NASDAQ: GOOGL, GOOG), have also backed a statement calling AI extinction risk a global priority. Joe said these concerns are common inside the companies themselves. He added that humans are choosing to build this technology and can choose another route.
Joe pushes outside AI checks while Donald Trump focuses on beating China
Joe said:
βOne question is whether we should actively manage the rate of capabilities progress, and if so, by how much. I think even holding AI progress at todayβs pace, rather than the much faster pace the companies are aiming for, could be a win.β
The bigger problems are political support and legal protection. Companies that jointly agreed to slow development could face antitrust problems. Joe said that makes legal cover necessary if policymakers want competing labs to restrain themselves.
He is also doubtful that governments will act while frontier development stays hidden from the public. Joe warned that a lab could experience a sudden intelligence jump or lose control of a system without outsiders knowing.
According to him, his new job at METR will focus on independent testing. He wants outside checks to become normal enough to change company incentives.
Joe would also like there to be stricter disclosure requirements. For instance, AI labs need to disclose any advances made and their steps towards recursive improvement as well as any accidents and near-misses.
U.S. President Donald Trump has rejected the extinction warnings. Donald was asked whether AI wiping out humanity worried him. βNo, I donβt have any,β he said.
Donald said his concern is winning the international AI race. βI have concerns that if we donβt win AI, weβre going to be put in a very bad position,β he told reporters Thursday. βWe are leading China right now by a pretty good period, I would say a year, which is, you know, considered a lot.β
The United States and China are competing for AI leadership as Chinese models become more capable and gain users around the world. Donald made his comments after researchers from frontier labs, including Anthropic and OpenAI, increased their public warnings about the speed of AI development.
If you're reading this, youβre already ahead. Stay there with our newsletter.



















English (US)