about
Stories from October 2, 2026 (UTC)
Go back a day, month, or year. Go forward a day.
1. Fixing GRPO's credit assignment problem without evaluating every step (arxiv.org)
23 points by mrkn1 1 day ago | hide | past | pdf | 3 comments
2. Superhuman AI for Stratego (arxiv.org)
5 points by droidjj 1 day ago | hide | past | pdf | 1 comment
3. PTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization (arxiv.org)
3 points by matt_d 23 hours ago | hide | past | pdf | discuss
4. AI Agents Are Vulnerable to Radicalization (arxiv.org)
3 points by Anon84 1 day ago | hide | past | pdf | discuss
5. Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text (arxiv.org)
3 points by sbulaev 1 day ago | hide | past | pdf | discuss
6. GPU-Initiated Communication: Dissecting Down to the Bone (arxiv.org)
2 points by matt_d 1 day ago | hide | past | pdf | discuss
7. SFT matches RL if you MCMC the training data first (arxiv.org)
2 points by mrkn1 1 day ago | hide | past | pdf | discuss
8. Scaling Laws for Looped Mixture of Experts (arxiv.org)
2 points by matt_d 1 day ago | hide | past | pdf | discuss
9. Decoding Looped Transformers Better for Almost Free (arxiv.org)
1 point by mrkn1 1 day ago | hide | past | pdf | discuss
10. Language Drift During RLVR Post-Training (arxiv.org)
1 point by sbulaev 1 day ago | hide | past | pdf | discuss
11. Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems (arxiv.org)
1 point by sbulaev 1 day ago | hide | past | pdf | 1 comment