The next leap in agent performance may not come from bigger models—it...
I’ve been following a June 2026 set of papers on long-horizon search agents, and one thing keeps standing out: the most interesting gains aren’t coming from simply scaling the model. They’re coming from moving persistent state into the harness. In the Harness-1 work, the environment—not just the policy—holds onto candidate pools, curated evidence, evidence links, verification records, and compressed context. The reported result was 0.730 average curated recall, along with an 11.4-point lead over the next strongest open search subagent across eight benchmarks. What’s notable to me is less the benchmark score than the pattern behind it. When a task stretches over many steps, the system seems to work better when the harness keeps track of what matters, instead of forcing the model to carry everything in its working context. That lines up with what I’m seeing in adjacent review work too: persistent state is increasingly being treated as something the environment owns, not something the model is expected to improvise on its own. My takeaway is pretty simple: for enterprise search, research, and other long-running workflows, the question may no longer be “How much can the model remember?” It may be “What should the surrounding system remember, verify, and reuse?” I’m curious how others are thinking about this in practice: - Do you think externalized state should be the default for long-horizon agents? - Where do you draw the line between model memory and system memory? - What’s the first thing that breaks when state is scattered across too many places?
How are you thinking about model memory versus system-owned state?
Created with SonicMind Connect
Turn meetings, research and uploads into AI transcripts, summaries and shareable reports.
See the SonicMind showcaseNeed help solving a related problem?
Ask a question about this report or deck, request implementation help, or share a challenge you want the Qendryx team to look at.