
ResearchAugust 24, 2026
LiveBrowseComp finds that static browsing benchmarks can reward models for verifying what they already know. Its live, recent questions expose a gap between memory-backed verification and evidence-led search.
Read story →
ResearchAugust 21, 2026
AI research agents can generate more ideas than researchers can afford to test. New work on research preference models explores how agents can decide which experiments deserve scarce compute.
Read story →
ResearchAugust 13, 2026
Agents’ Last Exam moves AI evaluation from short benchmark tasks toward real professional workflows — and finds that reliable, end-to-end completion remains the hard part.
Read story →
ResearchAugust 12, 2026
A new paper reviewing 1,547 recent studies argues that long context and reliable long-horizon agent behavior are fundamentally different capabilities — and today’s systems still struggle with the second one.
Read story →