| 2881. |
Emergence of Diffusion Models from Associative Memory (arxiv.org) |
|
9 points by fzliu on Jun 20, 2025 | hide | past | pdf | 2 comments
|
| 2882. |
Approximating Language Model Training Data from Weights (arxiv.org) |
|
2 points by jxmorris12 on Jun 20, 2025 | hide | past | pdf | discuss
|
| 2883. |
Interpreting Agent Behaviors in RL-Based Cyber-Battle Simulation Platforms (arxiv.org) |
|
1 point by PaulHoule on Jun 20, 2025 | hide | past | pdf | discuss
|
| 2884. |
Sekai: A Video Dataset Towards World Exploration (arxiv.org) |
|
2 points by badmonster on Jun 20, 2025 | hide | past | pdf | 1 comment
|
| 2885. |
ProtoReasoning: Prototypes as the Foundation for Generalizable Reasoning in LLMs (arxiv.org) |
|
2 points by simonpure on Jun 20, 2025 | hide | past | pdf | discuss
|
| 2886. |
Learning-based density-equalizing map (arxiv.org) |
|
2 points by PaulHoule on Jun 20, 2025 | hide | past | pdf | discuss
|
| 2887. |
Robustly Improving LLM Fairness in Realistic Settings via Interpretability (arxiv.org) |
|
2 points by scribu on Jun 18, 2025 | hide | past | pdf | discuss
|
| 2888. |
S1: Simple Test-Time Scaling (arxiv.org) |
|
3 points by bicepjai on Jun 18, 2025 | hide | past | pdf | discuss
|
| 2889. |
Style over Substance: Distilled Language Models Reason via Stylistic Replication (arxiv.org) |
|
5 points by curtsmith on Jun 18, 2025 | hide | past | pdf | discuss
|
| 2890. |
Self-Supervised Contrastive Learning Approximates Supervised CL (arxiv.org) |
|
3 points by PaulHoule on Jun 18, 2025 | hide | past | pdf | discuss
|
| 2891. |
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions (arxiv.org) |
|
2 points by fzliu on Jun 18, 2025 | hide | past | pdf | discuss
|
| 2892. |
Reasoning by Superposition: A Perspective on Chain of Continuous Thought (arxiv.org) |
|
60 points by danielmorozoff on Jun 18, 2025 | hide | past | pdf | 1 comment
|
| 2893. |
SageAttention3: Microscaling FP4 Attention. 5x Speed up (arxiv.org) |
|
2 points by 0xjunhao on Jun 18, 2025 | hide | past | pdf | discuss
|
| 2894. |
Large Language Models – The Future of Fundamental Physics? (arxiv.org) |
|
1 point by matteocantiello on Jun 18, 2025 | hide | past | pdf | discuss
|
| 2895. |
Serving Large Language Models on Huawei CloudMatrix384 (arxiv.org) |
|
3 points by fspeech on Jun 17, 2025 | hide | past | pdf | discuss
|
| 2896. |
CURE: A Dataset for Clinical Understanding and Retrieval Evaluation (arxiv.org) |
|
1 point by fzliu on Jun 17, 2025 | hide | past | pdf | discuss
|
| 2897. |
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions (arxiv.org) |
|
5 points by sjb326 on Jun 17, 2025 | hide | past | pdf | 2 comments
|
| 2898. |
Large Language Models and Emergence: A Complex Systems Perspective (arxiv.org) |
|
2 points by mathgenius on Jun 17, 2025 | hide | past | pdf | discuss
|
| 2899. |
Rethinking Text-Based Protein Understanding: Retrieval or LLM? (arxiv.org) |
|
1 point by PaulHoule on Jun 17, 2025 | hide | past | pdf | discuss
|
| 2900. |
Scaling On-Device GPU Inference for Large Generative Models (arxiv.org) |
|
5 points by Anon84 on Jun 17, 2025 | hide | past | pdf | discuss
|
| 2901. |
Is there a half-life for the success rates of AI agents? (arxiv.org) |
|
2 points by YeGoblynQueenne on Jun 17, 2025 | hide | past | pdf | discuss
|
| 2902. |
The Illusion of the Illusion of Thinking (arxiv.org) |
|
12 points by jedisct1 on Jun 16, 2025 | hide | past | pdf | 1 comment
|
| 2903. |
BanglaByT5: Byte-Level Modelling for Bangla (arxiv.org) |
|
2 points by PaulHoule on Jun 16, 2025 | hide | past | pdf | discuss
|
| 2904. |
How Do Olympiad Medalists Judge LLMs in Competitive Programming? (arxiv.org) |
|
3 points by npalli on Jun 16, 2025 | hide | past | pdf | 1 comment
|
| 2905. |
Improving Continual Pre-Training Through Seamless Data Packing (arxiv.org) |
|
2 points by PaulHoule on Jun 16, 2025 | hide | past | pdf | discuss
|
| 2906. |
Spherical CNNs (2018) (arxiv.org) |
|
22 points by rkp8000 on Jun 16, 2025 | hide | past | pdf | 3 comments
|
| 2907. |
Breaking Quadratic Barriers: A Non-Attention LLM for Ultra-Long Context Horizons (arxiv.org) |
|
70 points by PaulHoule on Jun 16, 2025 | hide | past | pdf | 23 comments
|
| 2908. |
Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging (arxiv.org) |
|
1 point by PaulHoule on Jun 16, 2025 | hide | past | pdf | discuss
|
| 2909. |
Extracting memorized pieces of books from open-weight language models (arxiv.org) |
|
109 points by fzliu on Jun 16, 2025 | hide | past | pdf | 109 comments
|
| 2910. |
Lossless Token Sequence Compression via Meta-Tokens (arxiv.org) |
|
2 points by PaulHoule on Jun 16, 2025 | hide | past | pdf | discuss
|
| More |