about
Stories from January 2, 2025 (UTC)
Go back a day, month, or year. Go forward a day.
1. TinyStories: How Small Can Language Models Be and Still Speak Coherent English? (2023) (arxiv.org)
218 points by tzury on Jan 2, 2025 | hide | past | pdf | 104 comments
2. Why transformers are obviously good models of language (arxiv.org)
6 points by jxmorris12 on Jan 2, 2025 | hide | past | pdf | discuss
3. Meta: Memory Layers at Scale (arxiv.org)
4 points by georgehill on Jan 2, 2025 | hide | past | pdf | discuss
4. The Overthinking of O1-Like LLMs (arxiv.org)
3 points by omarsar on Jan 2, 2025 | hide | past | pdf | 1 comment
5. MVQ: Efficient DNN Compression and Acceleration with Masked Vector Quantization (arxiv.org)
2 points by PaulHoule on Jan 2, 2025 | hide | past | pdf | discuss
6. Generative Modeling with Explicit Memory (arxiv.org)
2 points by PaulHoule on Jan 2, 2025 | hide | past | pdf | discuss
7. Scaling of Search and Learning (arxiv.org)
2 points by jonbaer on Jan 2, 2025 | hide | past | pdf | discuss
8. Reinforcement Learning for Multi-Intersection Traffic Signal Control (arxiv.org)
1 point by PaulHoule on Jan 2, 2025 | hide | past | pdf | discuss