
Terminal-Bench-Science evaluates AI agents on workflows drawn from researchers' own work. Scientists set the bar for scientific capability in AI rather than model developers or data vendors.
The benchmark is led by researchers at Stanford University and built by the Terminal-Bench team in collaboration with domain experts from multiple scientific disciplines and research institutions worldwide. It measures AI agent capabilities through a diverse set of challenging tasks.
Key Takeaways
Center
Not enough sources yet on this side — a blindspot summary will appear once outlets cover the story.
Not enough sources yet on this side — a blindspot summary will appear once outlets cover the story.
Get every side of the week's biggest story in your inbox.
Outcomes will appear once reporting identifies clear next steps.
Sets independent scientific standards for AI agent evaluation.
No attributed perspective yet — we only show quotes grounded in analysis.
Evaluated across 1 reporting source (0% Left · 100% Center · 0% Right)
Historical editorial baseline for Hacker News (XX)
Multi-factor accuracy index for Hacker News (XX) and related desks based on verifiable sourcing and editorial standards.
Factuality: High 100%
Ownership: Hacker News (XX) (Corporate)
Public conversation related to this story
Loading comments…