about
Show HN: A 6M-token movable window on a single 46GB GPU (arxiv.org)
7 points by Wetime 67 days ago | hide | past | pdf | 16 comments on HN

In plain words: A frozen language model stores solutions that passed an independent check, then answers new problems from those families by exact lookup instead of generating words. Four models each got 180 of 180 right with zero generated tokens; empty the memory and they solve nothing.

Abstract · A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever

Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take the opposite path: the model stays frozen, and a persistent memory of verified solutions grows beside it. Once a problem family is solved and has passed an independent verification step that never consults the answer key, every new instance of that family is answered at zero generation tokens, bit-exact, deterministically. Across 180 fresh instances spanning nine problem families, four architectures from four vendors - dense and mixture-of-experts - each score 180/180 at zero generation tokens per answer: execution-bound capability decoupled from parameter scaling. A negative control attributes the capability fully to the memory: emptied, it solves nothing. The same verify-before-store contract holds for open-ended reasoning: 88/88 consistency-gated acceptances across all four models, machine-checked formal proof, and reasoning-method transfer at 77/80. Memory selection takes 1.4 microseconds; a full reuse completes in 6-23 ms at 36 mWh. Approximate similarity retrieval selects the wrong item 94.3% of the time on a 4,500-item verified store where exact addressing makes zero errors. The store also serves as working context at a scale no shipped engine matches: a 6,000,000-token movable window on a single 46 GB GPU at flat memory, where vLLM stops at 30,399 tokens and SGLang silently truncates past 32,000. On published benchmarks, frontier models remain far ahead of any 12B at raw from-scratch reasoning; on everything this system has solved and verified, the comparison inverts: a frontier API call pays a fresh generation pass on every query, forever, while verified reuse costs zero tokens and returns the identical bits every time. A public testbench with free, rate-limited access accompanies this report: https://corbenic-galahad-bench.hf.space

Sietse Schelpe
arXiv:2607.23806 · cs.CL, cs.AI, cs.IR, cs.LG, cs.PF · submitted Jul 26, 2026
abstract · pdf · Industry experience report. 14 pages, 8 figures. Public testbench: https://corbenic-galahad-bench.hf.space; companion repository with SHA-256 provenance manifest: https://github.com/corbenicai/galahad-bench

add comment on HN

"No implementation detail, algorithm, or configuration is contained in this document by design."

+ odd page cuts, it's as-if no human has ever looked at this before uploading it.

These LLMs are absolute poison for some folks.
You couldn't be bothered to write a coherent summary of what this actually is and what it does, you just let the AI write some random noise, eh?
This paper show a new tool called Galahad.Normally, AI has to think and guess the answer every time, which costs time and money.

The knowledge of the model grows next to it not the model itself and no it is not the same as cache

No fine-tuning needed

It gives the exact same right answer every time, costs zero extra tokens, and saves lots of energy

Maybe it all works, but the paper is not trivial to decipher and the GitHub repository does not seem to exist. It doesn't seem to define what are the inputs to the system (what is a query? UTF-8 text? tokens?) and what are the outputs. It'd really help if the algorithm was written out step by step with all the expected type information included.

At first I thought it was similar to something I've built before as a long-term slowly degrading cache for augmenting an FFN by caching well-learned answers, answering by performing a beam search in the key space resulting in located key accuracy measure (how well it corresponds to the input query) and answer confidence (has it been a long time since verification?), but that's not quite it? It feels similar in some way, but is it?

So, you don't know what it is, either.
We handeld llm like a human brain we decopelled knowledge from the memory and build a memory layer that makes redoing things free and fast, so the llm can once it learned something solves it for free the next time
this sounds like ur explaining caching
Sounds like they're explaining magic. Because current models cannot learn and they cannot remember.
there is RL for LLMs which actually changes the weights but its more specialization than learning and wont counter the probabalistic nature of the thing
Training is not something you can just bolt on, and it generally requires even larger hardware than inference for a given model, and a huge amount of time. If you're aiming for "free", a OP claims, RL aint it.
how is it not the same as a cache it its exact description matches the description of a cache?
A cache remembers answers (only useful for the exact same question again). We remember the proven method and redo the work on every new question, so it solves ones it's never seen, which a cache simply can't.
100% generated. I skimmed the paper, and came out with a feeling of still not knowing what this is about.
We handeld llm like a human brain we decoupled knowledge from the memory and build a memory layer that makes redoing things free and fast, so the llm can once it learned something solves it for free the next time
Only a single 46Gb GPU? Wow AI sure is amazing tech.