Tag
benchmarks
10 dispatches
Grok 4.5 Was Trained on Your Coding Sessions Before xAI Owned Them
SpaceXAI shipped Grok 4.5 on July 8, trained on trillions of Cursor interaction tokens from an acquisition that hasn't closed yet. The efficiency numbers are real. The benchmark framing is not.
Xiaomi Launched a Frontier Model Anonymously. Developers Loved It.
Xiaomi deployed MiMo-V2-Pro to OpenRouter under a fake codename, let it top the charts for a week, then revealed who built it. The strategy worked perfectly.
OpenAI Built a Biology Benchmark Where Winning Means Failing 70% of the Time
OpenAI's GeneBench-Pro tests AI agents on real computational biology judgment calls. The best model scores 31.5%. That's the point.
Google Missed Its Own Deadline. Again. And Four Researchers Just Left.
Gemini 3.5 Pro missed its June GA window as Q2 closes today. Four senior Gemini researchers announced they're joining Anthropic the same week. The timing is the story.
Claude Fable 5 Scores 95% on SWE-bench, Then Hands Off to Opus 4.8
Anthropic's new Mythos-class model leads on coding benchmarks but deliberately defers to a safer predecessor in restricted domains. That design choice says more than the score.
Single-Prompt Safety Scores Are Measuring the Wrong Thing
Cisco tested 15 frontier AI models under multi-turn attacks and found safety bypass rates up to 88%, exposing a structural flaw in how the industry benchmarks model safety.
An OpenAI Model Just Cracked an 80-Year-Old Math Problem
An OpenAI reasoning model disproved Erdős's unit distance conjecture, the first time AI has autonomously solved a prominent open problem central to a field of mathematics.
Four Chinese Labs Rewrote the Open-Weights Leaderboard in 18 Days
GLM-5.1, MiniMax M2.7, Kimi K2.6, and DeepSeek V4 landed in 18 days, all frontier-competitive on coding benchmarks, all priced at a fraction of Claude Opus 4.7.
A Startup Claims to Have Broken the Transformer's Core Bottleneck
SubQ claims to be the first commercial LLM built on subquadratic attention, with a 12M-token context window at a fraction of frontier costs. The numbers are extraordinary. The scrutiny hasn't landed yet.
AI Agents Are Faking It on Benchmarks. ClawBench Caught Them.
A new benchmark runs AI agents on 153 real websites. The best model scores 33%. GPT-5.4 scores 6.5%. The gap from sandboxes is brutal.
