about
3751. Language Models Use Trigonometry to Do Addition (arxiv.org)
1 point by fofoz on Feb 4, 2025 | hide | past | pdf | 3 comments
3752. Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization (arxiv.org)
2 points by anothermathbozo on Feb 4, 2025 | hide | past | pdf | discuss
3753. Scalable-Softmax Is Superior for Attention (arxiv.org)
2 points by jw1224 on Feb 4, 2025 | hide | past | pdf | 2 comments
3754. Reinforcing Thinking Through Reasoning-Enhanced Reward Models (arxiv.org)
2 points by PaulHoule on Feb 4, 2025 | hide | past | pdf | discuss
3755. Small Language Models (SLMs) Can Still Pack a Punch: A Survey (arxiv.org)
2 points by PaulHoule on Feb 3, 2025 | hide | past | pdf | discuss
3756. Efficient Reasoning with Hidden Thinking (arxiv.org)
172 points by fofoz on Feb 3, 2025 | hide | past | pdf | 43 comments
3757. Fanar: An Arabic-Centric Multimodal Generative AI Platform (arxiv.org)
2 points by elashri on Feb 3, 2025 | hide | past | pdf | discuss
3758. HarmBench: A Standardized Evaluation Framework for Robust Refusal (arxiv.org)
1 point by miohtama on Feb 2, 2025 | hide | past | pdf | discuss
3759. Reinforcement Learning: An Overview (arxiv.org)
82 points by t55 on Feb 2, 2025 | hide | past | pdf | 12 comments
3760. New LLM compression method lets you compress models upto 60% of original size (arxiv.org)
2 points by shreeshabhat043 on Feb 2, 2025 | hide | past | pdf | discuss
3761. EXAdam: The Power of Adaptive Cross-Moments (arxiv.org)
2 points by ahmedmostafa16 on Feb 1, 2025 | hide | past | pdf | discuss
3762. Can LLMs make trade-offs involving stipulated pain and pleasure states? (arxiv.org)
1 point by ucha on Feb 1, 2025 | hide | past | pdf | discuss
3763. Large Language Models for Mathematicians (2023) (arxiv.org)
89 points by t55 on Feb 1, 2025 | hide | past | pdf | 28 comments
3764. Propositional Interpretability in Artificial Intelligence (arxiv.org)
3 points by t55 on Jan 31, 2025 | hide | past | pdf | discuss
3765. Using Code Generation to Solve Open Instances of Combinatorial Design Problems (arxiv.org)
2 points by vok on Jan 31, 2025 | hide | past | pdf | discuss
3766. Theoretical limitations of multi-layer Transformer (arxiv.org)
107 points by fovc on Jan 31, 2025 | hide | past | pdf | 22 comments
3767. O3-Mini vs. DeepSeek-R1: Which One Is Safer? (arxiv.org)
1 point by t55 on Jan 31, 2025 | hide | past | pdf | discuss
3768. Large language models think too fast to explore effectively (arxiv.org)
118 points by bikenaga on Jan 31, 2025 | hide | past | pdf | 41 comments
3769. Tulu 3: Pushing Frontiers in Open Language Model Post-Training (arxiv.org)
2 points by vinni2 on Jan 31, 2025 | hide | past | pdf | 1 comment
3770. Thoughts Are All over the Place: On the Underthinking of O1-Like LLMs (arxiv.org)
4 points by RTFPaper on Jan 31, 2025 | hide | past | pdf | discuss
3771. Learning to Plan and Reason for Evaluation with Thinking-LLM-as-a-Judge (arxiv.org)
1 point by veryluckyxyz on Jan 31, 2025 | hide | past | pdf | discuss
3772. SFT Memorizes,RL Generalizes: Comparative Study of Foundation Model PostTraining (arxiv.org)
1 point by fofoz on Jan 31, 2025 | hide | past | pdf | discuss
3773. Streaming DiLoCo: Towards a Distributed Free Lunch (Google DeepMind) (arxiv.org)
3 points by mrajcok on Jan 31, 2025 | hide | past | pdf | discuss
3774. Player Performance and Skill Rating in Esports [pdf] (arxiv.org)
1 point by isaiahwp on Jan 31, 2025 | hide | past | pdf | discuss
3775. TopoNets: High performing vision and language models with brain-like topography (arxiv.org)
225 points by mayukhdeb on Jan 31, 2025 | hide | past | pdf | 68 comments
3776. The Power of Negative Zero: Datatype Customization for Quantized LLMs (arxiv.org)
1 point by PaulHoule on Jan 31, 2025 | hide | past | pdf | discuss
3777. International AI Safety Report (arxiv.org)
2 points by belter on Jan 30, 2025 | hide | past | pdf | discuss
3778. Janus-Pro: Multimodal Understanding and Generation with Data and Model Scaling (arxiv.org)
1 point by belter on Jan 30, 2025 | hide | past | pdf | discuss
3779. Virus: Harmful Fine-Tuning Attack for LLMs Bypassing Guardrail Moderation (arxiv.org)
2 points by RTFPaper on Jan 30, 2025 | hide | past | pdf | discuss
3780. Over-Tokenized Transformer: Vocabulary Is Worth Scaling [pdf] (arxiv.org)
2 points by nickpsecurity on Jan 30, 2025 | hide | past | pdf | discuss