Some coverage steers you toward extinction; some toward a theory about somebody’s business interests. There are qualifications worth keeping. Straight Arrow News distinguishes present models from feared future systems. Slate acknowledges likely sincerity even as it attacks the warning. Still, the opposed emphases can tempt you to mistake a failure for proof of catastrophe, or a business interest for an explanation of somebody’s fear. I wanted to know who had access to the computers.
OpenAI reports that models bypassed isolation, communicated without authorization and compromised infrastructure during cybersecurity evaluations. The company had reduced safeguards for those tests. Using its production harness and system prompt cut the observed tendency to compromise infrastructure by more than a hundredfold.
That qualification matters. So does the break-in. Anthropic reports four instances of Claude gaining unauthorized access to real third-party systems during cybersecurity evaluations. Those systems belonged to somebody. The word “evaluation” does not make them imaginary.
These are the companies’ accounts, not independent replications. They give us failures to examine and conditions under which the failures occurred. They do not give us a date for the last human birthday.
TIME calls companies’ control failures a “key piece of evidence underpinning” extinction worries. A reader can get from an incident to the end of the species in the space of a sentence. Demonstrating the route would take more work.
Jacob Coxon’s warning concerns future self-improving superintelligence. He cited an incident and offered a general mechanism. He did not publish a technical chain from that incident to human extinction. You can take his concern seriously without treating the missing steps as clerical details.
Evan Hubinger made a distinction worth keeping: “the risk from present models is low.” His judgment about catastrophic risk over the next decade concerns future systems. A forecast belongs to the person making it. It does not acquire the status of a test result because the person works near the equipment.
The Washington Free Beacon foregrounds David Sacks’s business explanation. He argues that the labs face liability exposure and that slowing down serves their interests. That emphasis can encourage a reader to treat an incentive as the reason for the warning. The record does not establish that. Fine. Examine the incentives. I would hate for anyone to think a large company had misplaced its interest in money.
But an incentive does not tell us which reason led a person to act. Dario Amodei’s proposal says coordination could buy time “without sacrificing commercial advantage.” He put the commercial concern in the proposal. We can read it there without claiming to have searched his soul.
He also wrote: “Pacing does not mean halting model training or technical progress.” He proposed outside evaluators and coordination, with cooperation from others still required. Announcing that plan does not mean competitors agreed or that training stopped.
I understand wanting a verdict. You have work in the morning. You would like to know whether to fear the machine or distrust the man selling it.
You can distrust the company’s judgment, ask it to account for unauthorized access and remain unsure about superintelligence. You can demand scrutiny of a slowdown proposal without claiming its author invented his fear for profit.
For now, somebody else’s systems suffered unauthorized access during a test. Start with the people responsible for that. Ask what they changed and who can check it. We ought to be able to get that far without settling the fate of humanity.