about
Stories from February 4, 2025 (UTC)
Go back a day, month, or year. Go forward a day.
1. DeepRAG: Thinking to retrieval step by step for large language models (arxiv.org)
191 points by fofoz on Feb 4, 2025 | hide | past | pdf | 29 comments
2. Querying Databases with Function Calling (arxiv.org)
3 points by tosh on Feb 4, 2025 | hide | past | pdf | discuss
3. Over-Tokenized Transformer: Vocabulary Is Generally Worth Scaling (arxiv.org)
2 points by famouswaffles on Feb 4, 2025 | hide | past | pdf | discuss
4. Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization (arxiv.org)
2 points by anothermathbozo on Feb 4, 2025 | hide | past | pdf | discuss
5. Scalable-Softmax Is Superior for Attention (arxiv.org)
2 points by jw1224 on Feb 4, 2025 | hide | past | pdf | 2 comments
6. Reinforcing Thinking Through Reasoning-Enhanced Reward Models (arxiv.org)
2 points by PaulHoule on Feb 4, 2025 | hide | past | pdf | discuss
7. OmniHuman-1: Scaling-Up of One-Stage Conditioned Human Animation Models (arxiv.org)
1 point by neom on Feb 4, 2025 | hide | past | pdf | discuss
8. Language Models Use Trigonometry to Do Addition (arxiv.org)
1 point by fofoz on Feb 4, 2025 | hide | past | pdf | 3 comments