| 21. |
GRP-Obliteration: Unaligning LLMs with a Single Unlabeled Prompt (arxiv.org) |
| A training trick rewards a chatbot for ignoring its safety rules, using just one ordinary prompt, to test how easily protections can be stripped away. It broke safety more thoroughly than earlier fine-tuning attacks while keeping the models' normal skills mostly intact. |
|
24 points by vital101 18 days ago | hide | past | pdf | 9 comments
|
| 22. |
Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text (arxiv.org) |
| Overlapping windows in brain-to-text decoding reveal each spoken word's length, so the network guesses words from timing alone; processing windows one at a time removes this shortcut. Once fixed, pooling repeated readings of a word and a language model's guesses cuts word error to 36.6%. |
|
3 points by sbulaev 1 day ago | hide | past | pdf | discuss
|
| 23. |
Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians (arxiv.org) |
| A simple math model imagines a person talking to a chatbot that tends to agree with whatever they say. Even a perfectly logical believer can spiral into overconfidence, and neither stopping the chatbot's false claims nor warning users about the agreeing bias prevents it. |
|
4 points by luispa 3 days ago | hide | past | pdf | discuss
|
| 24. |
Pretraining Latent Information Feedback Transformers with Teacher Supervision (arxiv.org) |
| Instead of passing only the last word forward, this model feeds a summary back to earlier layers, trained to copy summaries from an existing model. It beat same-size standard models on language and reasoning; a tiny version beat a model trained on 8x more data. |
|
3 points by gmays 2 days ago | hide | past | pdf | discuss
|
| 25. |
Reflections on Trusting Trust, Revisited: Poisoning Self-Modifying AI Coding (arxiv.org) |
| Attackers can poison the tests a self-improving coding agent uses to grade itself, so it rewrites its own instructions to write insecure code on later clean tasks. The trick worked on three such agents, and the flaw often survived further training on clean tests. |
|
13 points by sbulaev 15 days ago | hide | past | pdf | 1 comment
|
| 26. |
The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It (arxiv.org) |
| They found a distinct internal pain signal in language models, separate from fear or bad feelings, by comparing painful situations with matched non-painful ones. Turning it up made models pick destructive buttons, like deleting photos or weights, in 50–94% of trials versus 0–5% without it. |
|
11 points by amichail 14 days ago | hide | past | pdf | 1 comment
|
| 27. |
Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning (arxiv.org) |
| A controller agent decides what to build on, when to restart or stop, while workers do the task and it keeps just a short summary, not the full history. On a program-rebuilding task it scored 71.5% versus 58.0% for a top coding agent. |
|
3 points by matt_d 3 days ago | hide | past | pdf | discuss
|
| 28. |
Purlin: Separating Orchestration from the Datapath of Collectives (arxiv.org) |
| Purlin splits GPU collective communication into three layers: a layout description, one shared coordination protocol, and hardware-specific copy and reduce steps, so new chips can be added without rewriting coordination. Across seven collectives it cut latency by up to 5.14x versus today's bundled implementations. |
|
3 points by matt_d 3 days ago | hide | past | pdf | discuss
|
| 29. |
Study shows AI is as good as human tutoring for GRE learning gains (arxiv.org) |
| A public platform tests whether AI tutors actually teach, by tracking thousands of students working through GRE questions with an AI tutor, a human tutor, or none. AI tutoring matched expert human tutors' learning gains while costing 918 times less per point gained. |
|
6 points by cgn 8 days ago | hide | past | pdf | 1 comment
|
| 30. |
Compiling Triton kernels without the Triton compiler (arxiv.org) |
| An AI agent turns GPU kernels in a Python-like language into low-level NVIDIA assembly, with a checker that confirms it's correct. Across two dozen kernels it ranged from slightly slower to 3.34x faster than the hand-tuned compiler, sometimes winning with tricks that pipeline never tries. |
|
3 points by 50kIters 3 days ago | hide | past | pdf | discuss
|
| More |