about
Stories from March 15, 2024 (UTC)
Go back a day, month, or year. Go forward a day.
1. Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking (arxiv.org)
280 points by hackerlight on Mar 15, 2024 | hide | past | pdf | 264 comments
2. MM1: Methods, Analysis and Insights from Multimodal LLM Pre-Training (arxiv.org)
3 points by kmdupree on Mar 15, 2024 | hide | past | pdf | discuss
3. Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient LMs (arxiv.org)
1 point by panabee on Mar 15, 2024 | hide | past | pdf | discuss
4. Wukong: Towards a Scaling Law for Large-Scale Recommendation (arxiv.org)
1 point by PaulHoule on Mar 15, 2024 | hide | past | pdf | discuss
5. Simple and Scalable Strategies to Continually Pre-Train Large Language Models (arxiv.org)
1 point by Anon84 on Mar 15, 2024 | hide | past | pdf | discuss
6. Apple MM1: Methods, Analysis and Insights from Multimodal LLM Pre-Training (arxiv.org)
1 point by CharlesW on Mar 15, 2024 | hide | past | pdf | 1 comment