about
Agentic Context Management: Memory and Cost as Architecture Problems (arxiv.org)
79 points by gdad 39 days ago | hide | past | pdf | 28 comments on HN

In plain words: Instead of storing and searching text, this treats an agent's memory as a lifecycle: deciding what to keep and shrinking context to fit a budget. A working system scored 92% on a long-term memory test, while piling up history makes token costs grow quadratically.

Abstract · Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems

Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs. Agents drown in their own accumulating history while paying a token cost that grows every turn, producing missing recalls within and across conversations. The incumbent response treats this as a storage-and-retrieval problem. We argue that framing is too narrow. Actively managing what an agent holds in mind is a lifecycle, not merely a store: it spans deciding what to remember, extracting and structuring it, choosing the right store per data type, consolidating and forgetting while preserving provenance, deciding what is relevant now, anticipating what is needed next, and compacting context to a budget without losing what matters. In serious production this operates not over a single user but across an organizational scope hierarchy. We name this discipline Agentic Context Management (ACM) and decompose it into five primitives: architecting, ingesting, scoping, anticipating, and compacting & consolidation. We then make the economic case: naive context accumulation grows token cost quadratically in conversation length, crude summarization buys linear cost at the price of an accuracy cliff, and only validated compaction achieves linear cost with preserved fidelity. We describe a reference implementation, Maximem Synap, that realizes the five primitives as a multi-tenant service and reports 92% on LongMemEval and 93.2% on LoCoMo under the configuration detailed in Section 6. We close with dimensions existing benchmarks do not yet capture, latency, token efficiency, and context-rot resistance, and the frontier of decision-level and organization-level context the category points toward.

Gaurav Dadhich
arXiv:2607.21503 · cs.AI, cs.IR · submitted Jul 23, 2026
abstract · pdf · html · 23 pages, 6 figures, 4 tables. Evaluation harness and study data: github.com/maximem-ai

add comment on HN
Also discussed: Jul 2026 (2 points, 0 comments)

Context pollution and rot are probably more important than memory, because facts can usually be retrieved if the agent is good at following breadcrumbs.

What's also the biggest killer is code rot. Agents are particularly good at death by thousand cuts. They implement something poorly, or incorrectly, or introduce a bad pattern into the project. Then they continue to amplify that badness over time, as they continue to copy from it on subsequent work. It spreads like a virus.

Keeping these seeds out of the project is very difficult, and cleaning up the rot is very difficult. It also seems like a hard problem to solve because following the existing codebase is something that is good when the code is good, but bad when it is bad. So, seemingly, the solution means more thinking and evaluation for every change that is being made.

> They implement something poorly, or incorrectly, or introduce a bad pattern into the project. Then they continue to amplify that badness over time, as they continue to copy from it on subsequent work. It spreads like a virus.

> Keeping these seeds out of the project is very difficult, and cleaning up the rot is very difficult.

My "aha" moment was when I realized this goes for all spheres of life where this tech is/will be introduced.

It goes for all spheres of life, full stop. I’m not sure if agents struggle with this because they learned it from humans, or if they struggle with it because it’s a universally challenging problem, but it’s something we share with them.
You full-stopped too soon: It may happen everywhere, but it doesn't happen the same way everywhere or for the same reasons.

LLMs will create different kinds of corruption than humans, because the underlying mechanisms are different. Our ability to manage type of corruption will depend on how whether they can be predicted by math or intuition.

You’re kind of begging the question that I asked after the full stop. I was wondering how much the underlying mechanism matters if it’s modelling the same environment to a similar level of effectiveness.
The comment highlighted how LLMs exacerbate the issue by entrenching the preexisting issues.
Indeed, and reply to comment seemed a "yes and" -- As with humans.

It's curious how much of these could apply to either:

https://en.wikipedia.org/wiki/Reconstructive_memory

https://en.wikipedia.org/wiki/Misinformation_effect

And many mechanisms exacerbate issues by entrenching preexisting issues.

I don't understand. The comment said if humans want to change the route, this tech makes it more difficult. Human inertia is X, inertia with this tech is X ^ Y. The Y is the issue being discussed.
"Out of the crooked timber of humanity, no straight thing was ever wrought"
> Then they continue to amplify that badness over time

Also, with "self-bias", models are also likely to grow new content into spots that match their subtle fingerprints from the past.

That might come at the expense of whatever corrected "we should avoid that and do this instead" alternative some human added for future architecture.

Yes so regular human and agentic evaluation of the coding agent output, scoring it on specific criteria?

https://github.com/harness/harness-evals

Truly. Doing this for coding agents is an interesting and different shaped problem.
ACM, that's the term that I'd been looking for - and your paper explains it clearly. At the end, most of LLM problems are context problems. Getting the correct knowledge into its context window without overpopulating it is the actual engineering effort for most agents. And the solution you present seems promising.

Both compaction with validation and predictive fetching are the way to go.

I do not want to write an implementation for this myself, and if Synap is that implementation, I'd like to ask you a few questions: 1. Does it work with context that's not just agent conversations, but rather documents? 2. Is it better than RAG on large dataset? 3. What does on-prem options look like?

Thanks Samyakk!

1. Yes, works on docs, agent conversations, human-conversations from different sources (Slack, JIRA, etc.). We have connectors for some of these as well; so it is plug and play 2. conventional RAG recall accuracy is quite low (50-60%) and latency is pretty high (seconds). But worst is the precision; you end up context stuffing to get acceptable recall 3. We do offer on-prem deployments, but only on sizeable annual contracts

I like to start with memory engineering then reach full system then reducing costs. This allows unlocking full potential of agents.
Interesting. Where can I read more about this?
I came up with this after building a few systems.

For me, if I put costs in my architecture, I'm limited heavily and that can easily change how memory is shaped dramatically. The opposite, putting memory in architecture, is not true. Memory engineering first, then full scale in the system then costs considerations.

In addition, I believe this is future friendly. Because AI is advancing and getting smarter and cheaper everyday.

I couldn't find a guide on this so I share my basic thoughts.

I've found that a simple markdown knowledgebase (with some useful extensions like semantic search & git context) is all I need to improve the memory of my agents. Even my non-coding agents have a memory repo.

Here's my implementation: https://hraness.com/kb

Context drift on retries is easily the most annoying part of this setup. Locking down the tool payload schema first was the only thing that worked for us
Im wondering how silent information loss is detected later and what exactly the validation score measures.
your website isnt working
Apologies. Should be working now.
"I"
Nice, is there any harness that implements this approach?
Ive never read a paper cover to cover before but after wrestling with opus 5s english this paper is such a relief to read, its like my eyes has been washed off opus stink
Haha! I am going to put this one up as a win! Thanks for reading! Hope you found it useful.
Yet, it is full of AI slop one-liners like: "The contest ahead is not over who stores the most data; it is over who manages context the best".