| 81. |
Jev-as-a-Judge: Accept When Confident, Escalate When Unsure (arxiv.org) |
| A cheap judge gives only the odds for each verdict, and its confidence decides whether to accept it or hand the task to a slower reasoning judge. On 1,610 held-out pairs this routing beat GPT-6 by 0.9 points at 41% of its cost. |
|
2 points by nico 4 days ago | hide | past | pdf | discuss
|
| 82. |
Schedules Are Solvable Symbols: Tuning-Free Compilation of Tile Programs (arxiv.org) |
| Loom is a compiler that turns chip layout and data movement into a puzzle solved at compile time, picking placement, communication, and tile sizes instead of hand-tuned kernels. On two chip generations it matched or beat the vendor's kernels on three AI operations without tuning. |
|
2 points by matt_d 4 days ago | hide | past | pdf | discuss
|
| 83. |
Shockingly Simple Self-retrospection Improves Agentic Models Without RL (arxiv.org) |
| After each coding attempt, the agent writes an explanation of what worked or failed, and is trained only on those explanation words—no rewards or outside teacher. It solved 49.2% of held-out coding tasks versus 48.0% for reward-based training, using half the updates. |
|
2 points by Betelbuddy 4 days ago | hide | past | pdf | discuss
|
| 84. |
Language Models Act on Hidden Valence (arxiv.org) |
| Tagging one of two meaningless zones with a 'good' or 'bad' internal brain pattern, then switching it off, showed which zone the model preferred. Models picked the good zone even when every visible word was identical, and worked to erase bad states. |
|
2 points by sbulaev 4 days ago | hide | past | pdf | discuss
|
| 85. |
CORVUS: Context Optimization and Reduction via Underlying Synchronization (arxiv.org) |
| Instead of pasting a fixed copy of every file it reads into its history, the agent keeps a live list of files and pulls in their current contents at each step. This cut input tokens by 9-50% while solving the same share of tasks. |
|
2 points by matt_d 5 days ago | hide | past | pdf | discuss
|
| 86. |
An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries (arxiv.org) |
| A chatbot saves the math from reading a prompt so follow-ups come fast; today each cloud keeps its copy. This vision shares those saved chunks across clouds, deciding where to store or recompute them by network speed and price to cut delay and cost. |
|
2 points by matt_d 5 days ago | hide | past | pdf | discuss
|
| 87. |
Recursive Self-Improvement via On-Policy Distillation for Reasoning (arxiv.org) |
| A model teaches itself using a copy that sees the answer as a guide, and both copies keep improving each round so new gains feed back in. On an 8-billion-parameter model this hit 65.97% on math contests, beating the frozen-guide version by 35.62 points. |
|
2 points by simonpure 5 days ago | hide | past | pdf | discuss
|
| 88. |
A-Mem: Agentic Memory for LLM Agents (arxiv.org) |
| A memory system that turns each new experience into a note with keywords and tags, then links it to related past notes and updates those notes as understanding grows. Tested on six different AI models, it beat the best existing memory systems at helping agents use past experience. |
|
2 points by jerlendds 5 days ago | hide | past | pdf | discuss
|
| 89. |
Open-Source E2E FHE Implementation for Privacy-Preserving Llama 3 8B Inference (arxiv.org) |
| Fully homomorphic encryption lets a server run a model on scrambled data it never sees; this system repacks data so weights and results need less re-encoding. It runs Llama 3 8B on encrypted input on one GPU, 4.51 times faster than the earlier approach. |
|
2 points by simonpure 6 days ago | hide | past | pdf | discuss
|
| 90. |
What and Whose Knowledge? Measuring Epistemic Diversity in Large Language Models (arxiv.org) |
| A study counted the distinct real-world claims 27 chatbots produced across 155 topics in 12 countries over three years, comparing them with web search results. Variety has grown over time but still trails search, and English knowledge crowds out local-language knowledge. |
|
2 points by daniel_iversen 7 days ago | hide | past | pdf | discuss
|
| More |