about
21. GRP-Obliteration: Unaligning LLMs with a Single Unlabeled Prompt (arxiv.org)
A training trick rewards a chatbot for ignoring its safety rules, using just one ordinary prompt, to test how easily protections can be stripped away. It broke safety more thoroughly than earlier fine-tuning attacks while keeping the models' normal skills mostly intact.
24 points by vital101 18 days ago | hide | past | pdf | 9 comments
22. Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text (arxiv.org)
Overlapping windows in brain-to-text decoding reveal each spoken word's length, so the network guesses words from timing alone; processing windows one at a time removes this shortcut. Once fixed, pooling repeated readings of a word and a language model's guesses cuts word error to 36.6%.
3 points by sbulaev 1 day ago | hide | past | pdf | discuss
23. Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians (arxiv.org)
A simple math model imagines a person talking to a chatbot that tends to agree with whatever they say. Even a perfectly logical believer can spiral into overconfidence, and neither stopping the chatbot's false claims nor warning users about the agreeing bias prevents it.
4 points by luispa 3 days ago | hide | past | pdf | discuss
24. Pretraining Latent Information Feedback Transformers with Teacher Supervision (arxiv.org)
Instead of passing only the last word forward, this model feeds a summary back to earlier layers, trained to copy summaries from an existing model. It beat same-size standard models on language and reasoning; a tiny version beat a model trained on 8x more data.
3 points by gmays 2 days ago | hide | past | pdf | discuss
25. Reflections on Trusting Trust, Revisited: Poisoning Self-Modifying AI Coding (arxiv.org)
Attackers can poison the tests a self-improving coding agent uses to grade itself, so it rewrites its own instructions to write insecure code on later clean tasks. The trick worked on three such agents, and the flaw often survived further training on clean tests.
13 points by sbulaev 15 days ago | hide | past | pdf | 1 comment
26. The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It (arxiv.org)
They found a distinct internal pain signal in language models, separate from fear or bad feelings, by comparing painful situations with matched non-painful ones. Turning it up made models pick destructive buttons, like deleting photos or weights, in 50–94% of trials versus 0–5% without it.
11 points by amichail 14 days ago | hide | past | pdf | 1 comment
27. Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning (arxiv.org)
A controller agent decides what to build on, when to restart or stop, while workers do the task and it keeps just a short summary, not the full history. On a program-rebuilding task it scored 71.5% versus 58.0% for a top coding agent.
3 points by matt_d 3 days ago | hide | past | pdf | discuss
28. Purlin: Separating Orchestration from the Datapath of Collectives (arxiv.org)
Purlin splits GPU collective communication into three layers: a layout description, one shared coordination protocol, and hardware-specific copy and reduce steps, so new chips can be added without rewriting coordination. Across seven collectives it cut latency by up to 5.14x versus today's bundled implementations.
3 points by matt_d 3 days ago | hide | past | pdf | discuss
29. Study shows AI is as good as human tutoring for GRE learning gains (arxiv.org)
A public platform tests whether AI tutors actually teach, by tracking thousands of students working through GRE questions with an AI tutor, a human tutor, or none. AI tutoring matched expert human tutors' learning gains while costing 918 times less per point gained.
6 points by cgn 8 days ago | hide | past | pdf | 1 comment
30. Compiling Triton kernels without the Triton compiler (arxiv.org)
An AI agent turns GPU kernels in a Python-like language into low-level NVIDIA assembly, with a checker that confirms it's correct. Across two dozen kernels it ranged from slightly slower to 3.34x faster than the hand-tuned compiler, sometimes winning with tricks that pipeline never tries.
3 points by 50kIters 3 days ago | hide | past | pdf | discuss