
The ‘Horizon Gap’ Explains Why Long Context Still Doesn’t Mean Long-Task Reliability
A new paper reviewing 1,547 recent studies argues that long context and reliable long-horizon agent behavior are fundamentally different capabilities — and today’s systems still struggle with the second one.
Read story