about
Stories from February 15, 2025 (UTC)
Go back a day, month, or year. Go forward a day.
1. Can We Trust AI Benchmarks? A Review of Current Issues in AI Evaluation (arxiv.org)
23 points by rntn on Feb 15, 2025 | hide | past | pdf | 4 comments
2. The hierarchy in HNSW is not necessary in high dimensions (arxiv.org)
5 points by blaise-muhirwa on Feb 15, 2025 | hide | past | pdf | 1 comment
3. Direct Ascent Synthesis: Hidden Generative Capabilities in Discriminative Models (arxiv.org)
2 points by kynez on Feb 15, 2025 | hide | past | pdf | discuss
4. What makes math problems hard for reinforcement learning: a case study (arxiv.org)
2 points by rntn on Feb 15, 2025 | hide | past | pdf | discuss
5. The Danger of Overthinking: The Reasoning-Action Dilemma in Agentic Tasks (arxiv.org)
2 points by stared on Feb 15, 2025 | hide | past | pdf | discuss
6. Geometry of Prompting: Distinct Mechanisms of Task Adaptation in Language Models (arxiv.org)
2 points by PlatinumSoup on Feb 15, 2025 | hide | past | pdf | discuss
7. Competitive Programming with Large Reasoning Models (arxiv.org)
1 point by xdavidliu on Feb 15, 2025 | hide | past | pdf | discuss