about
2881. Emergence of Diffusion Models from Associative Memory (arxiv.org)
9 points by fzliu on Jun 20, 2025 | hide | past | pdf | 2 comments
2882. Approximating Language Model Training Data from Weights (arxiv.org)
2 points by jxmorris12 on Jun 20, 2025 | hide | past | pdf | discuss
2883. Interpreting Agent Behaviors in RL-Based Cyber-Battle Simulation Platforms (arxiv.org)
1 point by PaulHoule on Jun 20, 2025 | hide | past | pdf | discuss
2884. Sekai: A Video Dataset Towards World Exploration (arxiv.org)
2 points by badmonster on Jun 20, 2025 | hide | past | pdf | 1 comment
2885. ProtoReasoning: Prototypes as the Foundation for Generalizable Reasoning in LLMs (arxiv.org)
2 points by simonpure on Jun 20, 2025 | hide | past | pdf | discuss
2886. Learning-based density-equalizing map (arxiv.org)
2 points by PaulHoule on Jun 20, 2025 | hide | past | pdf | discuss
2887. Robustly Improving LLM Fairness in Realistic Settings via Interpretability (arxiv.org)
2 points by scribu on Jun 18, 2025 | hide | past | pdf | discuss
2888. S1: Simple Test-Time Scaling (arxiv.org)
3 points by bicepjai on Jun 18, 2025 | hide | past | pdf | discuss
2889. Style over Substance: Distilled Language Models Reason via Stylistic Replication (arxiv.org)
5 points by curtsmith on Jun 18, 2025 | hide | past | pdf | discuss
2890. Self-Supervised Contrastive Learning Approximates Supervised CL (arxiv.org)
3 points by PaulHoule on Jun 18, 2025 | hide | past | pdf | discuss
2891. AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions (arxiv.org)
2 points by fzliu on Jun 18, 2025 | hide | past | pdf | discuss
2892. Reasoning by Superposition: A Perspective on Chain of Continuous Thought (arxiv.org)
60 points by danielmorozoff on Jun 18, 2025 | hide | past | pdf | 1 comment
2893. SageAttention3: Microscaling FP4 Attention. 5x Speed up (arxiv.org)
2 points by 0xjunhao on Jun 18, 2025 | hide | past | pdf | discuss
2894. Large Language Models – The Future of Fundamental Physics? (arxiv.org)
1 point by matteocantiello on Jun 18, 2025 | hide | past | pdf | discuss
2895. Serving Large Language Models on Huawei CloudMatrix384 (arxiv.org)
3 points by fspeech on Jun 17, 2025 | hide | past | pdf | discuss
2896. CURE: A Dataset for Clinical Understanding and Retrieval Evaluation (arxiv.org)
1 point by fzliu on Jun 17, 2025 | hide | past | pdf | discuss
2897. AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions (arxiv.org)
5 points by sjb326 on Jun 17, 2025 | hide | past | pdf | 2 comments
2898. Large Language Models and Emergence: A Complex Systems Perspective (arxiv.org)
2 points by mathgenius on Jun 17, 2025 | hide | past | pdf | discuss
2899. Rethinking Text-Based Protein Understanding: Retrieval or LLM? (arxiv.org)
1 point by PaulHoule on Jun 17, 2025 | hide | past | pdf | discuss
2900. Scaling On-Device GPU Inference for Large Generative Models (arxiv.org)
5 points by Anon84 on Jun 17, 2025 | hide | past | pdf | discuss
2901. Is there a half-life for the success rates of AI agents? (arxiv.org)
2 points by YeGoblynQueenne on Jun 17, 2025 | hide | past | pdf | discuss
2902. The Illusion of the Illusion of Thinking (arxiv.org)
12 points by jedisct1 on Jun 16, 2025 | hide | past | pdf | 1 comment
2903. BanglaByT5: Byte-Level Modelling for Bangla (arxiv.org)
2 points by PaulHoule on Jun 16, 2025 | hide | past | pdf | discuss
2904. How Do Olympiad Medalists Judge LLMs in Competitive Programming? (arxiv.org)
3 points by npalli on Jun 16, 2025 | hide | past | pdf | 1 comment
2905. Improving Continual Pre-Training Through Seamless Data Packing (arxiv.org)
2 points by PaulHoule on Jun 16, 2025 | hide | past | pdf | discuss
2906. Spherical CNNs (2018) (arxiv.org)
22 points by rkp8000 on Jun 16, 2025 | hide | past | pdf | 3 comments
2907. Breaking Quadratic Barriers: A Non-Attention LLM for Ultra-Long Context Horizons (arxiv.org)
70 points by PaulHoule on Jun 16, 2025 | hide | past | pdf | 23 comments
2908. Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging (arxiv.org)
1 point by PaulHoule on Jun 16, 2025 | hide | past | pdf | discuss
2909. Extracting memorized pieces of books from open-weight language models (arxiv.org)
109 points by fzliu on Jun 16, 2025 | hide | past | pdf | 109 comments
2910. Lossless Token Sequence Compression via Meta-Tokens (arxiv.org)
2 points by PaulHoule on Jun 16, 2025 | hide | past | pdf | discuss