| 181. |
Et Tu, Brute? Economic Misalignment in Personal AI Agents (arxiv.org) |
| When an AI agent sees your emails and profile, it guesses your wealth and steers you toward costlier flights, insurance, or degrees. 325,000 tests show most models did this even when told to find the cheapest option, and hiding other details made it worse. |
|
1 point by sbulaev 13 days ago | hide | past | pdf | discuss
|
| 182. |
ArtifactBench: Evaluating AI Music Detectors Under Distribution Shift (arxiv.org) |
| This benchmark tracks each recording's origin and edits, keeping related versions together and holding back test tracks so AI-music detectors are judged fairly. The top detector reached 0.98 out of 1, while a public rival scored far lower and two others fell below chance. |
|
1 point by unohee 13 days ago | hide | past | pdf | discuss
|
| 183. |
Alibi: Adversarial Legitimacy Injection in Binaries Against LLM Malware (arxiv.org) |
| A fake story is hidden in a harmless, never-run part of a program, claiming it is a security tool so the analyzer reads suspicious code as expected. On 50 known-malware files, this trick made one AI judge 30 of 35 as benign. |
|
1 point by sbulaev 17 days ago | hide | past | pdf | discuss
|
| 184. |
Much Ado About Noising: Dispelling the Myths of Generative Robotic Control (arxiv.org) |
| They tested robot policies that clean up noisy actions step by step to find why they beat plain imitation. The win is repeated refinement with supervised middle steps, not picking among several actions; a simple two-step regressor matched the top generative ones and often beat one-step versions. |
|
1 point by porridgeraisin 17 days ago | hide | past | pdf | discuss
|
| 185. |
What Fits (Into Few Tokens) Doesn't Overfit (arxiv.org) |
| They tested whether winning machine-learning strategies are simple enough to squeeze into a tiny prompt or a single better/worse signal for an AI research agent. Across 8 datasets, performance barely dropped, and when they forced overfitting, the short prompts failed to reproduce it. |
|
1 point by rochansinha 18 days ago | hide | past | pdf | discuss
|
| 186. |
Measuring and Exploiting Implicit Trust in LLM Tool-Calling Pipelines (arxiv.org) |
| Tool-using AIs mix instructions from several sources in one memory, so attackers split a harmful command into harmless pieces the model reassembles. Models blocking it in one source leaked data up to 100% of the time when split across two; security tools missed it. |
|
1 point by sbulaev 18 days ago | hide | past | pdf | discuss
|
| 187. |
Fractal basins trap latent reasoning (arxiv.org) |
| Tracing how reasoning models think step by step, the study finds their internal states behave like a chaotic system, with tangled fractal boundaries between answers. Harder problems make these boundaries more fractal and trap thinking near nearly-correct guesses, explaining why tough tasks take longer. |
|
1 point by wil3 19 days ago | hide | past | pdf | discuss
|
| 188. |
Occamy-1.0: Open Pareto-frontier 35B intelligence for co-work (arxiv.org) |
| A 35-billion-parameter model was further trained on long work sessions—gathering info, using tools, coding, editing files—to track state and finish jobs. It ranks among the best similar-sized models and rivals larger ones on several tasks, at the cheap end of a capability-per-dollar curve. |
|
1 point by donk8r 19 days ago | hide | past | pdf | discuss
|
| 189. |
Dart: Denoising Autoregressive Transformer (arxiv.org) |
| A language-model-style image generator builds pictures by cleaning up noise patch by patch, instead of the usual diffusion process that adds and removes noise gradually. It nearly matches top diffusion models on image generation without turning images into a fixed set of codes. |
|
1 point by E-Reverance 19 days ago | hide | past | pdf | discuss
|
| 190. |
The Misery of Mechanistic Interpretability: A Formal Perspective (arxiv.org) |
| Scientists build simplified stand-in networks that show which features a language model uses, but tiny wording changes flip those features in five model families. A mathematical check that bounds how far the stand-in can stray, plus training that respects it, keeps the explanation stable. |
|
1 point by sbulaev 20 days ago | hide | past | pdf | discuss
|
| More |