| 3751. |
Language Models Use Trigonometry to Do Addition (arxiv.org) |
|
1 point by fofoz on Feb 4, 2025 | hide | past | pdf | 3 comments
|
| 3752. |
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization (arxiv.org) |
|
2 points by anothermathbozo on Feb 4, 2025 | hide | past | pdf | discuss
|
| 3753. |
Scalable-Softmax Is Superior for Attention (arxiv.org) |
|
2 points by jw1224 on Feb 4, 2025 | hide | past | pdf | 2 comments
|
| 3754. |
Reinforcing Thinking Through Reasoning-Enhanced Reward Models (arxiv.org) |
|
2 points by PaulHoule on Feb 4, 2025 | hide | past | pdf | discuss
|
| 3755. |
Small Language Models (SLMs) Can Still Pack a Punch: A Survey (arxiv.org) |
|
2 points by PaulHoule on Feb 3, 2025 | hide | past | pdf | discuss
|
| 3756. |
Efficient Reasoning with Hidden Thinking (arxiv.org) |
|
172 points by fofoz on Feb 3, 2025 | hide | past | pdf | 43 comments
|
| 3757. |
Fanar: An Arabic-Centric Multimodal Generative AI Platform (arxiv.org) |
|
2 points by elashri on Feb 3, 2025 | hide | past | pdf | discuss
|
| 3758. |
HarmBench: A Standardized Evaluation Framework for Robust Refusal (arxiv.org) |
|
1 point by miohtama on Feb 2, 2025 | hide | past | pdf | discuss
|
| 3759. |
Reinforcement Learning: An Overview (arxiv.org) |
|
82 points by t55 on Feb 2, 2025 | hide | past | pdf | 12 comments
|
| 3760. |
New LLM compression method lets you compress models upto 60% of original size (arxiv.org) |
|
2 points by shreeshabhat043 on Feb 2, 2025 | hide | past | pdf | discuss
|
| 3761. |
EXAdam: The Power of Adaptive Cross-Moments (arxiv.org) |
|
2 points by ahmedmostafa16 on Feb 1, 2025 | hide | past | pdf | discuss
|
| 3762. |
Can LLMs make trade-offs involving stipulated pain and pleasure states? (arxiv.org) |
|
1 point by ucha on Feb 1, 2025 | hide | past | pdf | discuss
|
| 3763. |
Large Language Models for Mathematicians (2023) (arxiv.org) |
|
89 points by t55 on Feb 1, 2025 | hide | past | pdf | 28 comments
|
| 3764. |
Propositional Interpretability in Artificial Intelligence (arxiv.org) |
|
3 points by t55 on Jan 31, 2025 | hide | past | pdf | discuss
|
| 3765. |
Using Code Generation to Solve Open Instances of Combinatorial Design Problems (arxiv.org) |
|
2 points by vok on Jan 31, 2025 | hide | past | pdf | discuss
|
| 3766. |
Theoretical limitations of multi-layer Transformer (arxiv.org) |
|
107 points by fovc on Jan 31, 2025 | hide | past | pdf | 22 comments
|
| 3767. |
O3-Mini vs. DeepSeek-R1: Which One Is Safer? (arxiv.org) |
|
1 point by t55 on Jan 31, 2025 | hide | past | pdf | discuss
|
| 3768. |
Large language models think too fast to explore effectively (arxiv.org) |
|
118 points by bikenaga on Jan 31, 2025 | hide | past | pdf | 41 comments
|
| 3769. |
Tulu 3: Pushing Frontiers in Open Language Model Post-Training (arxiv.org) |
|
2 points by vinni2 on Jan 31, 2025 | hide | past | pdf | 1 comment
|
| 3770. |
Thoughts Are All over the Place: On the Underthinking of O1-Like LLMs (arxiv.org) |
|
4 points by RTFPaper on Jan 31, 2025 | hide | past | pdf | discuss
|
| 3771. |
Learning to Plan and Reason for Evaluation with Thinking-LLM-as-a-Judge (arxiv.org) |
|
1 point by veryluckyxyz on Jan 31, 2025 | hide | past | pdf | discuss
|
| 3772. |
SFT Memorizes,RL Generalizes: Comparative Study of Foundation Model PostTraining (arxiv.org) |
|
1 point by fofoz on Jan 31, 2025 | hide | past | pdf | discuss
|
| 3773. |
Streaming DiLoCo: Towards a Distributed Free Lunch (Google DeepMind) (arxiv.org) |
|
3 points by mrajcok on Jan 31, 2025 | hide | past | pdf | discuss
|
| 3774. |
Player Performance and Skill Rating in Esports [pdf] (arxiv.org) |
|
1 point by isaiahwp on Jan 31, 2025 | hide | past | pdf | discuss
|
| 3775. |
TopoNets: High performing vision and language models with brain-like topography (arxiv.org) |
|
225 points by mayukhdeb on Jan 31, 2025 | hide | past | pdf | 68 comments
|
| 3776. |
The Power of Negative Zero: Datatype Customization for Quantized LLMs (arxiv.org) |
|
1 point by PaulHoule on Jan 31, 2025 | hide | past | pdf | discuss
|
| 3777. |
International AI Safety Report (arxiv.org) |
|
2 points by belter on Jan 30, 2025 | hide | past | pdf | discuss
|
| 3778. |
Janus-Pro: Multimodal Understanding and Generation with Data and Model Scaling (arxiv.org) |
|
1 point by belter on Jan 30, 2025 | hide | past | pdf | discuss
|
| 3779. |
Virus: Harmful Fine-Tuning Attack for LLMs Bypassing Guardrail Moderation (arxiv.org) |
|
2 points by RTFPaper on Jan 30, 2025 | hide | past | pdf | discuss
|
| 3780. |
Over-Tokenized Transformer: Vocabulary Is Worth Scaling [pdf] (arxiv.org) |
|
2 points by nickpsecurity on Jan 30, 2025 | hide | past | pdf | discuss
|
| More |