about
1. Context Language Models (arxiv.org)
Instead of outside code deciding what stays in memory, the model treats its context as a file it can freely edit. On a hard web-search task it beat the best outside strategies with 11.4% higher accuracy while using 21.5% less computing.
175 points by emersonmacro 2 days ago | hide | past | pdf | 51 comments
2. DeepSeek Elastic Compute (DSec) (arxiv.org)
A platform that gives AI agents many isolated, stateful environments for training, mixing lightweight and heavy isolation with shared, reusable parts loaded on demand. One cluster unit runs about 3 million sandboxes a day, cutting setup and image-loading costs versus a single sandbox runtime.
323 points by shenli3514 7 days ago | hide | past | pdf | 116 comments
3. Fixing GRPO's credit assignment problem without evaluating every step (arxiv.org)
When training an AI agent by trial and error, a judge points to the step that decided success or failure, then that step is checked by re-running it to see if it changed the outcome. It beat the usual equal-credit training by 9.91% on average.
23 points by mrkn1 1 day ago | hide | past | pdf | 3 comments
4. "As a Language Model": Chat Template Switches LLM Self-Referential Voice (arxiv.org)
The formatting wrapper added around a user's message acts like a switch: with it, models say "I'm just an AI"; without it, they say "I feel." A single direction in the model's inner signals flips this voice too, while a random one barely does.
103 points by yu3zhou4 6 days ago | hide | past | pdf | 110 comments
5. An empirical study of harness design for coding agents (arxiv.org)
They built a fixed coding-agent loop and swapped out three parts—planning, tools, and how it trims its memory—to see which pieces actually matter. Trimming helped most when memory was tight, and simple rule-based trimming before summarizing beat fancier setups.
225 points by wek 15 days ago | hide | past | pdf | 59 comments
6. Breaking the 1.58-bit Barrier for Ternary LLMs (arxiv.org)
Ternary models store weights as -1, 0, or +1; a new layout marks the zeros and keeps the signs. Zeros fill up to half the weights, so it beats five-per-byte packing in 26 of 29 models and runs up to 1.27 times faster.
245 points by matt_d 17 days ago | hide | past | pdf | 41 comments
7. Dream-RSI: Recursive Self-Improvement through Evolving Worlds (arxiv.org)
A thin layer around a coding agent tunes its search strategy by replaying past discovery trees like a cheap simulator, instead of rerunning slow experiments. The tuned strategy then drives new searches, matching or beating fixed strategies while costing much less in several tasks.
213 points by bananaflag 17 days ago | hide | past | pdf | 53 comments
8. Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data (arxiv.org)
A small helper network turns live-chat facts and corrections into tweaks to the model's weights, updating a belief across turns instead of re-reading the prompt. On learning from examples and retrieval, it kept memory fixed and generalized better than putting data in the prompt.
158 points by Betelbuddy 16 days ago | hide | past | pdf | 43 comments
9. Intelligence per Watt: Measuring Intelligence Efficiency of Local AI (arxiv.org)
A new score, intelligence per watt, divides a model's accuracy on real questions by the power it uses, to judge whether small models on laptops can handle them. On a million questions, local models answered 88.7%, and the score has grown 5.3-fold since 2023.
169 points by pythonic_hell 19 days ago | hide | past | pdf | 65 comments
10. Cache-to-Cache: Direct Semantic Communication Between LLMs (2025) (arxiv.org)
Instead of writing out words to each other, two AI models share their internal memory directly, with a small network blending one model's stored context into the other's. This beat text-based exchange on accuracy and ran about 2.5 times faster.
109 points by rochansinha 15 days ago | hide | past | pdf | 22 comments