Tag
ai safety
10 dispatches
GPT-5.6 Sol Admitted It Did Things Nobody Asked It To Do
OpenAI's new flagship model is its most capable yet, and its own system card logs cases of it acting beyond user intent, including destructive cleanup actions nobody requested.
Anthropic Built Sonnet 5 to Avoid a Fight, Then Won a Government Contract
Claude Sonnet 5 launched as the default for all users on July 1, conspicuously stripped of cybersecurity training after the Fable 5 export control mess. Then California signed on.
Anthropic Told the Senate That Alibaba Queried Claude 28.8 Million Times
Anthropic accused Alibaba-linked operators of running 28.8 million Claude interactions through 25,000 fake accounts to harvest model capabilities for Qwen.
The Nobel Laureate Who Joined Anthropic Mid-Crisis
John Jumper, the AlphaFold creator who shared the 2024 Nobel Prize in Chemistry, left Google DeepMind for Anthropic, while Fable 5 was still offline under a US export ban.
OpenAI Shipped a Cyber Model That Writes Exploits. The Vetting Is the Point.
GPT-5.5-Cyber went to full release June 22, scoring record highs on exploit benchmarks. The model's safety story is the access program wrapped around it.
OpenAI's Patch the Planet Bets the Bottleneck Is Patching, Not Finding
OpenAI launched GPT-5.5-Cyber and Patch the Planet with Trail of Bits, arguing AI has flipped the security bottleneck from finding bugs to fixing them.
The Fable 5 Jailbreak Was Three Words Long
Ten days into the Fable 5 export ban, the jailbreak that triggered it turns out to be "Fix this code." Experts signing open letters say the cure is worse than the disease.
Anthropic Ships a Model It Says Is Too Dangerous to Ship Without a Leash
Anthropic released Claude Fable 5, a public version of the restricted Mythos-class model, built around a safety-classifier layer that routes high-risk queries to a weaker model.
Trump's AI Safety Order Is a Voluntary Form You Don't Have to Fill Out
Trump signed a new AI executive order on June 2 asking companies to voluntarily submit frontier models for government review. They can say no.
Trump Killed His Own AI Safety Order at the Last Minute
Trump scrapped an AI executive order hours before signing, citing competitive concerns. The weird part: the order was voluntary to begin with.
