Tag
research
8 dispatches
OpenAI Built a Biology Benchmark Where Winning Means Failing 70% of the Time
OpenAI's GeneBench-Pro tests AI agents on real computational biology judgment calls. The best model scores 31.5%. That's the point.
The Nobel Laureate Who Joined Anthropic Mid-Crisis
John Jumper, the AlphaFold creator who shared the 2024 Nobel Prize in Chemistry, left Google DeepMind for Anthropic, while Fable 5 was still offline under a US export ban.
Google Lost the Transformer's Co-Author. Then AlphaFold. Same Week.
Noam Shazeer, co-author of "Attention Is All You Need," left Google for OpenAI on June 18. Days later, Nobel laureate John Jumper left for Anthropic. Both exits land at once.
Single-Prompt Safety Scores Are Measuring the Wrong Thing
Cisco tested 15 frontier AI models under multi-turn attacks and found safety bypass rates up to 88%, exposing a structural flaw in how the industry benchmarks model safety.
Karpathy Joined Anthropic to Train Claude Using Claude
Andrej Karpathy joined Anthropic's pretraining team in May 2026. The specific job: use Claude to accelerate the research that makes Claude better.
An OpenAI Model Just Cracked an 80-Year-Old Math Problem
An OpenAI reasoning model disproved Erdős's unit distance conjecture, the first time AI has autonomously solved a prominent open problem central to a field of mathematics.
A Startup Claims to Have Broken the Transformer's Core Bottleneck
SubQ claims to be the first commercial LLM built on subquadratic attention, with a 12M-token context window at a fraction of frontier costs. The numbers are extraordinary. The scrutiny hasn't landed yet.
AI Agents Are Faking It on Benchmarks. ClawBench Caught Them.
A new benchmark runs AI agents on 153 real websites. The best model scores 33%. GPT-5.4 scores 6.5%. The gap from sandboxes is brutal.
