about
Stories from April 1, 2025 (UTC)
Go back a day, month, or year. Go forward a day.
1. Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad (arxiv.org)
6 points by mauriziocalo on Apr 1, 2025 | hide | past | pdf | 1 comment
2. Large Language Models Are Unreliable for Cyber Threat Intelligence (arxiv.org)
5 points by belter on Apr 1, 2025 | hide | past | pdf | discuss
3. Research in AI for SWE is nowhere close to finished (arxiv.org)
4 points by minimario on Apr 1, 2025 | hide | past | pdf | discuss
4. Self-Vocabularizing Training for Neural Machine Translation (arxiv.org)
3 points by PaulHoule on Apr 1, 2025 | hide | past | pdf | discuss
5. Unlearning via Model Merging (arxiv.org)
3 points by kiyanwang on Apr 1, 2025 | hide | past | pdf | discuss
6. Benchmarking and Improving Robustness of Reward Models with Transformed Inputs (arxiv.org)
3 points by PaulHoule on Apr 1, 2025 | hide | past | pdf | discuss
7. Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad (arxiv.org)
3 points by oidar on Apr 1, 2025 | hide | past | pdf | 1 comment
8. SoK: Prompt Hacking of Large Language Models (2024) (arxiv.org)
2 points by hentrep on Apr 1, 2025 | hide | past | pdf | discuss
9. Large Language Models Share Representations of Latent Grammatical Concepts (arxiv.org)
2 points by Jimmc414 on Apr 1, 2025 | hide | past | pdf | discuss
10. What the Fuck Is Artificial General Intelligence? (arxiv.org)
1 point by mindcrime on Apr 1, 2025 | hide | past | pdf | discuss
11. Effectively Controlling Reasoning Models Through Thinking Intervention (arxiv.org)
1 point by belter on Apr 1, 2025 | hide | past | pdf | discuss