about
Stories from June 3, 2024 (UTC)
Go back a day, month, or year. Go forward a day.
1. Grokfast: Accelerated Grokking by Amplifying Slow Gradients (arxiv.org)
117 points by johnsutor on Jun 3, 2024 | hide | past | pdf | 39 comments
2. Evaluating Gemini Models for Dangerous Capabilities (arxiv.org)
3 points by fishfish on Jun 3, 2024 | hide | past | pdf | discuss
3. LLMs achieve adult human performance on higher-order theory of mind tasks (arxiv.org)
2 points by tosh on Jun 3, 2024 | hide | past | pdf | 1 comment
4. Leveraging Human Revisions for Improving Text-to-Layout Models (arxiv.org)
2 points by PaulHoule on Jun 3, 2024 | hide | past | pdf | discuss
5. There and Back Again: The AI Alignment Paradox (arxiv.org)
2 points by belter on Jun 3, 2024 | hide | past | pdf | discuss
6. Transformers are SSMs (Mamba-2) (arxiv.org)
2 points by jasondavies on Jun 3, 2024 | hide | past | pdf | discuss
7. Is Complexity an Illusion? (arxiv.org)
2 points by broyojo on Jun 3, 2024 | hide | past | pdf | 1 comment
8. Faithful Logical Reasoning via Symbolic Chain-of-Thought (arxiv.org)
2 points by burakemir on Jun 3, 2024 | hide | past | pdf | discuss
9. Swarm Parallelism: Training Large Models on Poorly Connected Devices (arxiv.org)
2 points by kwindla on Jun 3, 2024 | hide | past | pdf | discuss
10. Kotlin ML Pack: Technical Report (arxiv.org)
1 point by belter on Jun 3, 2024 | hide | past | pdf | discuss