about
3061. Outcome-Based Reinforcement Learning to Predict the Future (arxiv.org)
99 points by bturtel on May 27, 2025 | hide | past | pdf | 15 comments
3062. Learning to Reason Without External Rewards (arxiv.org)
4 points by epipolar on May 27, 2025 | hide | past | pdf | discuss
3063. Gradient-Based Program Repair: Fixing Bugs in Continuous Program Spaces (arxiv.org)
16 points by andre15silva on May 27, 2025 | hide | past | pdf | discuss
3064. Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents (arxiv.org)
2 points by pontiacbandit8 on May 27, 2025 | hide | past | pdf | discuss
3065. Grammars of Formal Uncertainty (arxiv.org)
34 points by barthelomew on May 27, 2025 | hide | past | pdf | 5 comments
3066. ARC-NCA: Towards Developmental Solutions to the Abstraction and Reasoning Corpus (arxiv.org)
1 point by jekude on May 27, 2025 | hide | past | pdf | discuss
3067. Effective Reinforcement Learning for Reasoning in Language Models (arxiv.org)
4 points by obastani on May 26, 2025 | hide | past | pdf | discuss
3068. Scaling RNNs to Billions of Parameters with Zero Order (arxiv.org)
7 points by fchaubard on May 26, 2025 | hide | past | pdf | 3 comments
3069. Better Zero-Shot Reasoning with Role-Play Prompting (arxiv.org)
2 points by virtual_rf on May 26, 2025 | hide | past | pdf | discuss
3070. The Paradox of Prompting: Less Detail Makes AI More Human (arxiv.org)
3 points by virtual_rf on May 26, 2025 | hide | past | pdf | discuss
3071. Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents (arxiv.org)
4 points by nobody9999 on May 26, 2025 | hide | past | pdf | 1 comment
3072. Large Language Model-Powered Agent for C to Rust Code Translation (arxiv.org)
2 points by MarcoDewey on May 25, 2025 | hide | past | pdf | discuss
3073. Interactive Post-Training for Vision-Language-Action Models (arxiv.org)
4 points by badmonster on May 25, 2025 | hide | past | pdf | discuss
3074. Neural Thermodynamic Laws for Large Language Model Training (arxiv.org)
6 points by anticensor on May 25, 2025 | hide | past | pdf | discuss
3075. Gen2seg: Generative Models Enable Generalizable Instance Segmentation (arxiv.org)
2 points by danielmorozoff on May 25, 2025 | hide | past | pdf | discuss
3076. When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs (arxiv.org)
3 points by Anon84 on May 24, 2025 | hide | past | pdf | discuss
3077. The Dangers of Browsing AI Agents (arxiv.org)
2 points by wslh on May 24, 2025 | hide | past | pdf | discuss
3078. An Invitation to Neuroalgebraic Geometry (arxiv.org)
2 points by IdealeZahlen on May 24, 2025 | hide | past | pdf | discuss
3079. A Consequentialist Critique of Binary Classification Evaluation Practices (arxiv.org)
3 points by alexmolas on May 24, 2025 | hide | past | pdf | discuss
3080. Beyond Semantics: Unreasonable Effectiveness of Reasonless Intermediate Tokens (arxiv.org)
138 points by nyrikki on May 23, 2025 | hide | past | pdf | 66 comments
3081. Understanding Generative AI Capabilities in Everyday Image Editing Tasks (arxiv.org)
5 points by taesiri on May 23, 2025 | hide | past | pdf | 1 comment
3082. Let Androids Dream of Electric Sheep: A Human-Like Image Implication Understand (arxiv.org)
1 point by badmonster on May 23, 2025 | hide | past | pdf | discuss
3083. SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA (arxiv.org)
2 points by tanelpoder on May 23, 2025 | hide | past | pdf | 1 comment
3084. The effectiveness of Large Language Models in the mechanical design domain (arxiv.org)
1 point by PaulHoule on May 22, 2025 | hide | past | pdf | discuss
3085. MMaDA: Multimodal Large Diffusion Language Models (arxiv.org)
2 points by pr337h4m on May 22, 2025 | hide | past | pdf | discuss
3086. People who use ChatGPT for writing are accurate detectors of AI-generated text (arxiv.org)
1 point by speckx on May 22, 2025 | hide | past | pdf | discuss
3087. Hierarchical-Chain-of-Generation for Complex Attributes Text-to-3D Generation (arxiv.org)
2 points by PaulHoule on May 22, 2025 | hide | past | pdf | discuss
3088. Bielik v3 Small: Technical Report (arxiv.org)
1 point by PaulHoule on May 22, 2025 | hide | past | pdf | discuss
3089. Project Sid: Many-Agent Simulations Toward AI Civilization (2024) (arxiv.org)
2 points by virtual_rf on May 22, 2025 | hide | past | pdf | 2 comments
3090. Beyond Semantics: Unreasonable Effectiveness of Reasonless Intermediate Tokens (arxiv.org)
2 points by NotInOurNames on May 22, 2025 | hide | past | pdf | discuss