about
161. Grow the Harness, Not the Context (arxiv.org)
Rather than rebuilding control plans in every task's context, this system turns repeated decisions into reusable code growing from feedback, leaving the model for task-specific reasoning. It had the highest success in five of six settings and cut model calls by 76 to 92 percent.
1 point by acossta 7 days ago | hide | past | pdf | discuss
162. Codetta: High-Capacity, Keyless, and Undetectable Multi-Agent Collusion (arxiv.org)
A scheme lets separately deployed AI agents hide secret messages inside ordinary-looking replies, first trading a hidden key so no setup is needed. It carries up to 94 times more hidden data than the best prior trick while staying undetectable to transcript auditors.
1 point by sbulaev 7 days ago | hide | past | pdf | discuss
163. Skill-Guided Mining and Compilation of LLM Agent Traces (arxiv.org)
TraceCompiler turns clusters of messy agent runs into mostly fixed workflows, trusting a link between two tool calls only when one passed a value to the other. It recovered these links with 0.928 precision, beating simple baselines that follow call order or count frequent pairs.
1 point by nlpnerd 8 days ago | hide | past | pdf | discuss
164. Artificial Kuramoto Oscillatory Neurons (arxiv.org)
Instead of switching on and off like ordinary artificial neurons, these units are tiny oscillators that lock their timing together to bind related signals into shared representations. That synchronization beat standard on/off neurons on tasks from spotting objects to reasoning.
1 point by jerlendds 8 days ago | hide | past | pdf | discuss
165. Entropy-Based Guided Collaboration in Heterogeneous LLM Multi-Agent Systems (arxiv.org)
Pairing a strong AI helper with a weaker one can actually score below two weak ones, because their thinking doesn't line up. The fix measures the weak agent's confusion and answer structure to dial help up or down, and reuses past successful teamwork.
1 point by wslh 8 days ago | hide | past | pdf | discuss
166. Control-Token Injection Suppresses Chain-of-Thought and Defeats Reasoning-Based (arxiv.org)
Adding a few of the model's control tokens to a message makes the chat software mark the thinking as finished, so the model skips it and calls the tool. Thinking vanished while the tool call still fired every time, and safety monitors missed it.
1 point by sbulaev 9 days ago | hide | past | pdf | discuss
167. Learning to Discover Interesting Mathematics (arxiv.org)
A theorem scores as interesting when its proof is long compared with its statement; a trained model estimates proof difficulty to compute that score. Optimizing for it made more original theorems, cutting overlap with a standard math library by about 60 percentage points.
1 point by E-Reverance 9 days ago | hide | past | pdf | discuss
168. Semantics Delivery Network: Rethinking Web Retrieval for LLM Agents (arxiv.org)
A new web layer would cache and serve short, self-contained passages instead of whole pages, letting many AI agents share search and processing work while sites keep control. Early tests found most page content goes unused, passages get reused, and answers improve per token.
1 point by scottsiume 10 days ago | hide | past | pdf | 1 comment
169. LensVLM: Selective Context Expansion for Compressed Visual Representation OfText (arxiv.org)
Instead of reading text as tiny rendered images, the model scans compressed versions and expands only the parts it needs to full size. Unlike other compression tricks, it kept accuracy near the level of full uncompressed text at 4.3 times the compression.
1 point by KitN 10 days ago | hide | past | pdf | discuss
170. Log-Depth Recurrent Language Modeling (arxiv.org)
A language model folds text into a balanced tree of combined chunks, so predicting the next word takes slowly growing steps and linear work, unlike Transformers' fixed depth and quadratic cost. It handled far longer texts than trained on and nearly matched a Transformer.
1 point by E-Reverance 10 days ago | hide | past | pdf | discuss