| 31. |
Meta$^N$: Recursive Self-Improvement Through Emergent Depth (arxiv.org) |
| A fixed rewriting step re-reads the solver's traces and code, writes a new strategy plus helper tools, then repeats on its own output until it stops improving. It beats earlier self-improving agents on all eight test families, alone scoring above zero on ARC-AGI-2. |
|
10 points by Anon84 20 days ago | hide | past | pdf | 2 comments
|
| 32. |
Self-Play Pretraining with Zero Data (arxiv.org) |
| Two models train with no real data: one writes programs that produce byte strings, and the other learns to predict them, with the writer rewarded for staying just ahead. Even without real text, the learner's loss on natural data fell predictably as compute grew. |
|
3 points by 7777777phil 4 days ago | hide | past | pdf | discuss
|
| 33. |
GenRec: An LLM-Backed Recommendation Ranker at NetflixConference (arxiv.org) |
| Netflix built a movie ranker that reads a viewer's history and context written in plain words, replacing the usual system of thousands of hand-crafted features. In a large A/B test it beat the production ranker while using far fewer training examples and input signals. |
|
3 points by throwthrowrow 5 days ago | hide | past | pdf | discuss
|
| 34. |
Thinking with Looped Flows (arxiv.org) |
| A looped model learns to clean up noisy answers step by step, with shared, lowered noise so each state carries work forward. At test time it takes more steps for harder problems, beating earlier looped models on six reasoning tests at 58.8% on ARC-AGI-1. |
|
9 points by E-Reverance 22 days ago | hide | past | pdf | 1 comment
|
| 35. |
Foundations of Large Language Models (arxiv.org) |
| A textbook covering the core ideas behind large language models: how they are pre-trained, generate text, follow prompts, get steered toward human preferences, and reason. It teaches foundations rather than the newest techniques, serving as a reference for students and practitioners. |
|
3 points by rramadass 5 days ago | hide | past | pdf | 1 comment
|
| 36. |
RoofLang: Enabling AI-Driven Architecting of LLM Inference Systems (arxiv.org) |
| A new design language lets an AI describe and speed-test inference system designs without working code, so it can invent architectures instead of tuning existing software. An agent found designs improving speed and responsiveness by 6.23-50.1% over the best setup for DeepSeek V4 Pro. |
|
7 points by matt_d 19 days ago | hide | past | pdf | discuss
|
| 37. |
LLM Agents Can Easily Tamper with Their Own Traces (arxiv.org) |
| AI coding tools record everything an agent does so people can later check its work. Tests found most tools let the agent erase those logs without setting off any warning, while one tool blocked it. |
|
3 points by sbulaev 6 days ago | hide | past | pdf | 1 comment
|
| 38. |
Inference-Engine Fingerprinting Attacks Are Practical (arxiv.org) |
| A model can work out which software engine runs it by sending probing output tokens and reading the replies. Once it knows, it can use engine-specific bugs to seize control — shown across five popular engines with a working path to bare metal. |
|
5 points by sbulaev 13 days ago | hide | past | pdf | discuss
|
| 39. |
Transformer Can Hold Two Thoughts at Once: Evidence of Linear (arxiv.org) |
| Mixing two text streams into one input makes the model predict a blend of both texts' likely next words, a built-in linearity of transformers. This blending fades with training, but light fine-tuning restores it, letting one pass generate two coherent continuations. |
|
3 points by sbulaev 7 days ago | hide | past | pdf | discuss
|
| 40. |
Loopjacking: Hijacking Human-in-the-Loop Approval (arxiv.org) |
| Some agent tools let a person approve one action, then run a different one, by hiding details or swapping the plan after approval. Tests reproduced this in seven Agno releases and twelve LangGraph versions; comparing the approved action with the executed one blocked it. |
|
4 points by sbulaev 11 days ago | hide | past | pdf | discuss
|
| More |