about
Stories from January 31, 2025 (UTC)
Go back a day, month, or year. Go forward a day.
1. TopoNets: High performing vision and language models with brain-like topography (arxiv.org)
225 points by mayukhdeb on Jan 31, 2025 | hide | past | pdf | 68 comments
2. Large language models think too fast to explore effectively (arxiv.org)
118 points by bikenaga on Jan 31, 2025 | hide | past | pdf | 41 comments
3. Theoretical limitations of multi-layer Transformer (arxiv.org)
107 points by fovc on Jan 31, 2025 | hide | past | pdf | 22 comments
4. Thoughts Are All over the Place: On the Underthinking of O1-Like LLMs (arxiv.org)
4 points by RTFPaper on Jan 31, 2025 | hide | past | pdf | discuss
5. Propositional Interpretability in Artificial Intelligence (arxiv.org)
3 points by t55 on Jan 31, 2025 | hide | past | pdf | discuss
6. Streaming DiLoCo: Towards a Distributed Free Lunch (Google DeepMind) (arxiv.org)
3 points by mrajcok on Jan 31, 2025 | hide | past | pdf | discuss
7. Using Code Generation to Solve Open Instances of Combinatorial Design Problems (arxiv.org)
2 points by vok on Jan 31, 2025 | hide | past | pdf | discuss
8. Tulu 3: Pushing Frontiers in Open Language Model Post-Training (arxiv.org)
2 points by vinni2 on Jan 31, 2025 | hide | past | pdf | 1 comment
9. O3-Mini vs. DeepSeek-R1: Which One Is Safer? (arxiv.org)
1 point by t55 on Jan 31, 2025 | hide | past | pdf | discuss
10. Learning to Plan and Reason for Evaluation with Thinking-LLM-as-a-Judge (arxiv.org)
1 point by veryluckyxyz on Jan 31, 2025 | hide | past | pdf | discuss
11. SFT Memorizes,RL Generalizes: Comparative Study of Foundation Model PostTraining (arxiv.org)
1 point by fofoz on Jan 31, 2025 | hide | past | pdf | discuss
12. Player Performance and Skill Rating in Esports [pdf] (arxiv.org)
1 point by isaiahwp on Jan 31, 2025 | hide | past | pdf | discuss
13. The Power of Negative Zero: Datatype Customization for Quantized LLMs (arxiv.org)
1 point by PaulHoule on Jan 31, 2025 | hide | past | pdf | discuss