|
|
Stories from April 1, 2025 (UTC)
|
| 1. |
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad (arxiv.org) |
|
6 points by mauriziocalo on Apr 1, 2025 | hide | past | pdf | 1 comment
|
| 2. |
Large Language Models Are Unreliable for Cyber Threat Intelligence (arxiv.org) |
|
5 points by belter on Apr 1, 2025 | hide | past | pdf | discuss
|
| 3. |
Research in AI for SWE is nowhere close to finished (arxiv.org) |
|
4 points by minimario on Apr 1, 2025 | hide | past | pdf | discuss
|
| 4. |
Self-Vocabularizing Training for Neural Machine Translation (arxiv.org) |
|
3 points by PaulHoule on Apr 1, 2025 | hide | past | pdf | discuss
|
| 5. |
Unlearning via Model Merging (arxiv.org) |
|
3 points by kiyanwang on Apr 1, 2025 | hide | past | pdf | discuss
|
| 6. |
Benchmarking and Improving Robustness of Reward Models with Transformed Inputs (arxiv.org) |
|
3 points by PaulHoule on Apr 1, 2025 | hide | past | pdf | discuss
|
| 7. |
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad (arxiv.org) |
|
3 points by oidar on Apr 1, 2025 | hide | past | pdf | 1 comment
|
| 8. |
SoK: Prompt Hacking of Large Language Models (2024) (arxiv.org) |
|
2 points by hentrep on Apr 1, 2025 | hide | past | pdf | discuss
|
| 9. |
Large Language Models Share Representations of Latent Grammatical Concepts (arxiv.org) |
|
2 points by Jimmc414 on Apr 1, 2025 | hide | past | pdf | discuss
|
| 10. |
What the Fuck Is Artificial General Intelligence? (arxiv.org) |
|
1 point by mindcrime on Apr 1, 2025 | hide | past | pdf | discuss
|
| 11. |
Effectively Controlling Reasoning Models Through Thinking Intervention (arxiv.org) |
|
1 point by belter on Apr 1, 2025 | hide | past | pdf | discuss
|
|