about
Stories from February 11, 2026 (UTC)
Go back a day, month, or year. Go forward a day.
1. Misconduct in Post-Selections and Deep Learning (2024) (arxiv.org)
3 points by bjourne 235 days ago | hide | past | pdf | discuss
2. GRP-Obliteration: Unaligning LLMs with a Single Unlabeled Prompt [pdf] (arxiv.org)
2 points by janandonly 235 days ago | hide | past | pdf | discuss
3. Attention Sinks and Compression Valleys in LLMs (arxiv.org)
1 point by alexkranias 235 days ago | hide | past | pdf | discuss
4. The Hot Mess of AI: How Does Misalignment Scale with Model Intelligence (arxiv.org)
1 point by schmuhblaster 235 days ago | hide | past | pdf | discuss
5. Harmless reward hacks generalize to shutdown evasion and dictatorship in GPT-4.1 (arxiv.org)
1 point by toliveistobuild 235 days ago | hide | past | pdf | 1 comment