Tag
ai development
2 dispatches
OpenAI Built a Biology Benchmark Where Winning Means Failing 70% of the Time
OpenAI's GeneBench-Pro tests AI agents on real computational biology judgment calls. The best model scores 31.5%. That's the point.
Karpathy Joined Anthropic to Train Claude Using Claude
Andrej Karpathy joined Anthropic's pretraining team in May 2026. The specific job: use Claude to accelerate the research that makes Claude better.
