| 1. |
Context Language Models (arxiv.org) |
| Instead of outside code deciding what stays in memory, the model treats its context as a file it can freely edit. On a hard web-search task it beat the best outside strategies with 11.4% higher accuracy while using 21.5% less computing. |
|
175 points by emersonmacro 2 days ago | hide | past | pdf | 51 comments
|
| 2. |
DeepSeek Elastic Compute (DSec) (arxiv.org) |
| A platform that gives AI agents many isolated, stateful environments for training, mixing lightweight and heavy isolation with shared, reusable parts loaded on demand. One cluster unit runs about 3 million sandboxes a day, cutting setup and image-loading costs versus a single sandbox runtime. |
|
323 points by shenli3514 7 days ago | hide | past | pdf | 116 comments
|
| 3. |
Fixing GRPO's credit assignment problem without evaluating every step (arxiv.org) |
| When training an AI agent by trial and error, a judge points to the step that decided success or failure, then that step is checked by re-running it to see if it changed the outcome. It beat the usual equal-credit training by 9.91% on average. |
|
23 points by mrkn1 1 day ago | hide | past | pdf | 3 comments
|
| 4. |
"As a Language Model": Chat Template Switches LLM Self-Referential Voice (arxiv.org) |
| The formatting wrapper added around a user's message acts like a switch: with it, models say "I'm just an AI"; without it, they say "I feel." A single direction in the model's inner signals flips this voice too, while a random one barely does. |
|
103 points by yu3zhou4 6 days ago | hide | past | pdf | 110 comments
|
| 5. |
An empirical study of harness design for coding agents (arxiv.org) |
| They built a fixed coding-agent loop and swapped out three parts—planning, tools, and how it trims its memory—to see which pieces actually matter. Trimming helped most when memory was tight, and simple rule-based trimming before summarizing beat fancier setups. |
|
225 points by wek 15 days ago | hide | past | pdf | 59 comments
|
| 6. |
Breaking the 1.58-bit Barrier for Ternary LLMs (arxiv.org) |
| Ternary models store weights as -1, 0, or +1; a new layout marks the zeros and keeps the signs. Zeros fill up to half the weights, so it beats five-per-byte packing in 26 of 29 models and runs up to 1.27 times faster. |
|
245 points by matt_d 17 days ago | hide | past | pdf | 41 comments
|
| 7. |
Dream-RSI: Recursive Self-Improvement through Evolving Worlds (arxiv.org) |
| A thin layer around a coding agent tunes its search strategy by replaying past discovery trees like a cheap simulator, instead of rerunning slow experiments. The tuned strategy then drives new searches, matching or beating fixed strategies while costing much less in several tasks. |
|
213 points by bananaflag 17 days ago | hide | past | pdf | 53 comments
|
| 8. |
Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data (arxiv.org) |
| A small helper network turns live-chat facts and corrections into tweaks to the model's weights, updating a belief across turns instead of re-reading the prompt. On learning from examples and retrieval, it kept memory fixed and generalized better than putting data in the prompt. |
|
158 points by Betelbuddy 16 days ago | hide | past | pdf | 43 comments
|
| 9. |
Intelligence per Watt: Measuring Intelligence Efficiency of Local AI (arxiv.org) |
| A new score, intelligence per watt, divides a model's accuracy on real questions by the power it uses, to judge whether small models on laptops can handle them. On a million questions, local models answered 88.7%, and the score has grown 5.3-fold since 2023. |
|
169 points by pythonic_hell 19 days ago | hide | past | pdf | 65 comments
|
| 10. |
Cache-to-Cache: Direct Semantic Communication Between LLMs (2025) (arxiv.org) |
| Instead of writing out words to each other, two AI models share their internal memory directly, with a small network blending one model's stored context into the other's. This beat text-based exchange on accuracy and ran about 2.5 times faster. |
|
109 points by rochansinha 15 days ago | hide | past | pdf | 22 comments
|
| More |