about
Zero-Mem: Zero-Token Memory Operations for LLM Agents (arxiv.org)
101 points by theanonymousone 59 days ago | hide | past | pdf | 12 comments on HN

In plain words: Instead of using an LLM to write and search summaries of past chats, it keeps the original traces and organizes them by linked entities and by time, then pulls from both to answer. It kept similar accuracy while cutting memory-operation time by 57.6%.

Abstract

LLM agents need memory to act consistently over long interactions, yet many systems use additional LLM calls to operate that memory. Generating intermediate records and mediating their retrieval adds recurring token and time costs, while omitted or merged details can obscure the original evidence. We ask whether structured memory access requires generation at all. Zero-Mem introduces \emph{zero-token memory operations}: no step outside final question answering invokes an LLM or consumes LLM input or output tokens; encoder computation is accounted for separately. Zero-Mem preserves original interaction traces as its source of record. It organizes the traces in two complementary ways. An entity--context graph exposes connections across interactions, while a temporal hierarchy preserves conversational locality and session state. For each query, Zero-Mem weighs the two views, retrieves from both, and follows their structure to recover supporting relations or surrounding context. Deterministic calibration first discards conflicting evidence and then keeps the reader's answer grounded in the retrieved traces. Only the final-QA reader invokes an LLM. Across long-memory and long-context question-answering benchmarks, Zero-Mem achieves competitive performance while eliminating LLM calls and LLM-token consumption from memory operations. With the same final-QA reader and context budget, it reduces memory-operation time cost by 57.6\% relative to the fastest compared baseline. Ablations support the contribution of the two views and their query-dependent coordination. Overall, the results show that structured agent memory need not generate an intermediate representation of the past. After peer review, the code and implementation details will be available at \textcolor{blue}{https://github.com/TheMoon0815/Zero-mem}.

Yilin Xiao, Zhehan Zhu, Yujing Zhang, Jin Chen, Zijin Hong, Luyao Zhuang, Qinggang Zhang, Shengyuan Chen, Xiaocao Ouyang, Lingfei Ren, Xiao Huang
arXiv:2607.29377 · cs.CL · submitted Jul 31, 2026
abstract · pdf · html

add comment on HN

I am working on the same thing right now. However, unlike storing conversations in an external retrieval system, I use a local LLM to store the conversation's KV cache and perform retrieval directly on that cache. The method involves running a prefill pass and, after obtaining the attention scores, filtering for the corpus segments that received attention.

This aligns with the "zero tokens" approach described in this paper. :)

I tested it on the LoCoMo used in this paper, and also LongMemEval, both achieved SOTA results.

I was doing something similar where I saved user input/model output in a multi-depth node style storage system (each depth having more precise details) with the focus on the model having accurate user fed information. I was mostly focused on retrieval of accurate / useful information based on user query (injecting the database node as additional, high confidence information)

Once this(Zero-mem) passes it's peer review, I may have to see if my system can handle something similar instead/in addition.

I'm quite excited to see growth in these different ways of eliminating token's.

Long winded aside, @langs, have you published your work on this?

https://github.com/AttemorySystem/attemory/ stars and issues are welcome :)

Using attention for retrieval was inspired by a comment I saw in here long time ago: Prediction and retrieval are two sides of the same coin; to predict better, you must retrieve more accurately.

I'm still working on the improvement of algorithms, my tests shows the performance and accuracy will be improved a lot in the next release.

What the difference of doing this vs semantic search over indexed fact chunks? This is RAG right?
Yes, it uses modern technology to tackle an old problem: retrieval.
"A local Qwen3.5 retrieval model attends over the indexed memory"

Hrm.

I am not an expert on LLMs, but what's preventing you from treating the context as 'virtual memory', and using the attention matrix to 'blank out' tokens which have very low weights and will not contribute much to the input? I imagine most tokens are like this, and you can skip computation on 95% of an 1M (or practically infinite) context.
Hah I built a similar thing, stored a few million tokens chunked and precomputed winth Qwen a3e (best ratio of kv-size to tokens after chunking).

Some custom kernels and I was able to find all the relevant paragraphs with full force of qwen reasoning within 0.3s, and with a summary round within 0.7s.

Downside - required 200GB ram/vram ;) A few GBs for model and most of it for caching kvs.

This is actually quite easy to implement at the harness level and the NER can be way more naive because of the typical nature of LLM dialogue (programming, long running tasks etc).
Can confirm. I did a clean-room implementation of this in Rust just by reading the paper. Will update once they release the reference implementation. https://github.com/ptaranat/zeromem
Care to elaborate?
The useful contribution is not zero token cost; it is removing generative rewriting from memory. Preserving original traces avoids a subtle auditability failure: once an LLM compresses an interaction, retrieval is grounded in the summary's omissions rather than the evidence.

I would still want a harder benchmark around mutation and contradiction. If an entity changes attributes across sessions, can the graph and temporal hierarchy preserve both states, surface the conflict, and show which trace justified the answer? The 57.6% time reduction is compelling, but for production agents I would measure unsupported-answer rate and evidence recall under stale, conflicting, and adversarial traces. Encoder compute and index-maintenance cost should also sit beside token cost; otherwise "zero-token" risks being read as "free."