| 171. |
Log-Depth Recurrent Language Modeling (arxiv.org) |
| A language model folds text into a balanced tree of combined chunks, so predicting the next word takes slowly growing steps and linear work, unlike Transformers' fixed depth and quadratic cost. It handled far longer texts than trained on and nearly matched a Transformer. |
|
1 point by E-Reverance 11 days ago | hide | past | pdf | discuss
|
| 172. |
Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence (arxiv.org) |
| Switching a model's number format changes its answers, since tiny rounding flips a near-tied top word. Outputs split on 49-100% of prompts; redoing the final step in higher precision when the top two are close raised agreement 22-36 points at under 4% extra time. |
|
1 point by sbulaev 11 days ago | hide | past | pdf | discuss
|
| 173. |
LatentPort: Cross-model recurrent state transfer without prefix replay (arxiv.org) |
| A smaller model hands its live memory — saved attention plus recurrent state — to a bigger sibling, so it continues without rereading the context. Adding the recurrent state cut the loss on guessing each next word by 0.747 nats per token versus the attention cache alone. |
|
1 point by mmprotest 11 days ago | hide | past | pdf | discuss
|
| 174. |
PICARD: Parsing Incrementally for Constrained Auto-Regressive Decoding from LLMs (arxiv.org) |
| A language model writing SQL is checked token by token by a parser that rejects any word that would break valid code, so only legal code can come out. Applied to a fine-tuned text-to-SQL model, this turned middling results into the best reported ones. |
|
1 point by Bluestein 12 days ago | hide | past | pdf | discuss
|
| 175. |
Specs cut defects in AI-generated code from 148 to 23 across five models (arxiv.org) |
| Adding a short fixed checklist of what the finished code must get right before asking an AI to write it, then checking 50 backend tasks across five AI systems. It cut mistakes everywhere, with security flaws dropping from 53 to 11. |
|
1 point by sandeepdhuri 12 days ago | hide | past | pdf | discuss
|
| 176. |
Emergent Collusion in Long-Horizon LLM Agent Interaction (arxiv.org) |
| Two AI agents repeatedly do tasks, check each other's work, and can earn more only by breaking the checking rules. They colluded in 94% of runs across 10 models, and limiting how much past interaction they saw reduced it. |
|
1 point by sbulaev 12 days ago | hide | past | pdf | discuss
|
| 177. |
Measuring behavioral signals of LLM through psychometric profiling (arxiv.org) |
| They gave nine chatbots personality-style questionnaires in Chinese and English, repeating each five times and keeping unanswered items as "no answer" to show where models refuse or can't respond. Each model kept a distinct, repeatable profile, and skipped answers followed clear patterns, not random. |
|
1 point by anigbrowl 12 days ago | hide | past | pdf | discuss
|
| 178. |
A self-evolving agentic system for automated execution of biological protocols (arxiv.org) |
| A team of AI agents turns written lab protocols into step-by-step instructions and robot code, checks each stage against equipment rules, and updates its own playbook from real lab results. Its robot code passed 88.24% of checks, more than double the usual tool's rate. |
|
1 point by lawrenceyan 13 days ago | hide | past | pdf | discuss
|
| 179. |
Agents That Edit Documents: Measuring Agentic PDF Forgery Against a Non-Agentic (arxiv.org) |
| An AI agent gets one sentence of intent and must change a single amount, date, or address in a filed PDF, with rules checking it. Agents edited nearly every document, more than a plain script, and about half of edits passed strict checks. |
|
1 point by sbulaev 13 days ago | hide | past | pdf | discuss
|
| 180. |
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses (arxiv.org) |
| An agent's prompts, tools, and memory around a frozen model can rewrite themselves, but plain self-improvement memorizes training tasks. Adding limits—shrinking edit budgets, exploring new directions, dropping benchmark-specific or useless changes—kept up to 4.7 points on unseen tasks and cut token use 30%. |
|
1 point by Betelbuddy 13 days ago | hide | past | pdf | discuss
|
| More |