# Beyond pass@1: A Reliability Science Framework for Long-Horizon LLM Agents

> Aggregate pass@1 drops 24 percentage points (76.3% to 52.1%) as task duration grows across 10 models and 23,392 episodes. Memory scaffolds never helped any model and hurt 6 of 10 - the standard long-horizon intervention is empirically wrong.

Canonical URL: https://kravhal.kcsatish.com/insights/week-29
Edition: Week 16
Tags: Agents, LLMs, Optimization
Reading time: 8 min read

---

This is a mirror of an article first published in the AI & Automation Chronicle.

Full text with the original formatting: https://chronicle.kcsatish.com/posts/week-29
Markdown of the original: https://chronicle.kcsatish.com/posts/week-29.md
Structured JSON of the original: https://chronicle.kcsatish.com/api/v1/posts/week-29.json

Cite the Chronicle as the publication of record for the research claims in this article.
