about
2551. NoLiMa: Long-Context Evaluation Beyond Literal Matching (arxiv.org)
2 points by fzliu on Aug 15, 2025 | hide | past | pdf | discuss
2552. Distillation Scaling Laws (arxiv.org)
5 points by brandonb on Aug 15, 2025 | hide | past | pdf | discuss
2553. D2F – We made dLLMs 2.5x faster than LLaMA3 (arxiv.org)
5 points by zengzihao on Aug 15, 2025 | hide | past | pdf | 3 comments
2554. Efficient Attention Mechanisms for Large Language Models: A Survey (arxiv.org)
2 points by PaulHoule on Aug 15, 2025 | hide | past | pdf | discuss
2555. Capabilities of GPT-5 on Multimodal Medical Reasoning (arxiv.org)
1 point by amrrs on Aug 14, 2025 | hide | past | pdf | discuss
2556. Dr. Boot: Bootstrapping Program Synthesis Language Models to Perform Repairing (arxiv.org)
2 points by PaulHoule on Aug 14, 2025 | hide | past | pdf | discuss
2557. Contrastive Approach for Smart Ponzi Scheme Detecter with More Negative Samples (arxiv.org)
1 point by PaulHoule on Aug 14, 2025 | hide | past | pdf | discuss
2558. AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs? (arxiv.org)
1 point by PaulHoule on Aug 14, 2025 | hide | past | pdf | discuss
2559. Mathematical Computation and Reasoning Errors by Large Language Models (arxiv.org)
2 points by badmonster on Aug 14, 2025 | hide | past | pdf | discuss
2560. Demystifying NCCL: An In-Depth Analysis of GPU Communication Protocols and Algos (arxiv.org)
1 point by charleshn on Aug 13, 2025 | hide | past | pdf | discuss
2561. FilBench: Can LLMs Understand and Generate Filipino Language? (arxiv.org)
2 points by bananatype on Aug 13, 2025 | hide | past | pdf | 1 comment
2562. Large Language Models Do Not Simulate Human Psychology (arxiv.org)
1 point by Anon84 on Aug 13, 2025 | hide | past | pdf | discuss
2563. Scaling Recommender Transformers to One Billion Parameters (arxiv.org)
1 point by PaulHoule on Aug 13, 2025 | hide | past | pdf | discuss
2564. A Comprehensive Survey of Self-Evolving AI Agents [pdf] (arxiv.org)
94 points by SerCe on Aug 13, 2025 | hide | past | pdf | 29 comments
2565. SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model (arxiv.org)
2 points by ra7 on Aug 12, 2025 | hide | past | pdf | discuss
2566. Capabilities of GPT-5 on Multimodal Medical Reasoning (arxiv.org)
5 points by Anon84 on Aug 12, 2025 | hide | past | pdf | discuss
2567. G-Core: A Simple, Scalable and Balanced RLHF Trainer (arxiv.org)
2 points by PaulHoule on Aug 12, 2025 | hide | past | pdf | discuss
2568. AI agents fail tasks 70% of the time (arxiv.org)
23 points by JTbane on Aug 12, 2025 | hide | past | pdf | 8 comments
2569. Training language models to be warm and empathetic makes them less reliable (arxiv.org)
358 points by Cynddl on Aug 12, 2025 | hide | past | pdf | 375 comments
2570. Tricks or Traps? A Deep Dive into RL for LLM Reasoning (arxiv.org)
2 points by elashri on Aug 12, 2025 | hide | past | pdf | discuss
2571. GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models [pdf] (arxiv.org)
417 points by SerCe on Aug 12, 2025 | hide | past | pdf | 83 comments
2572. GLM-4.5: Agentic, Reasoning, and Coding (Arc) Foundation Models (arxiv.org)
3 points by Anon84 on Aug 11, 2025 | hide | past | pdf | discuss
2573. An analytic theory of creativity in convolutional diffusion models (arxiv.org)
3 points by derbOac on Aug 11, 2025 | hide | past | pdf | discuss
2574. Improving Generative Ad Text on Facebook Using Reinforcement Learning (arxiv.org)
2 points by yorwba on Aug 11, 2025 | hide | past | pdf | discuss
2575. Gemini Robotics: Bringing AI into the Physical World (arxiv.org)
2 points by Anon84 on Aug 11, 2025 | hide | past | pdf | discuss
2576. A Principled Framework to Use Data Bias for OOD Generation (arxiv.org)
1 point by PaulHoule on Aug 11, 2025 | hide | past | pdf | discuss
2577. Tribe: TRImodal Brain Encoder for whole-brain fMRI response prediction (arxiv.org)
2 points by georgehill on Aug 11, 2025 | hide | past | pdf | discuss
2578. Modern Methods in Associative Memory (arxiv.org)
5 points by liamdgray on Aug 10, 2025 | hide | past | pdf | 1 comment
2579. Is Chain-of-Thought Reasoning of LLMs a Mirage? (arxiv.org)
3 points by jerlendds on Aug 10, 2025 | hide | past | pdf | discuss
2580. Design Patterns for Securing LLM Agents Against Prompt Injections (arxiv.org)
2 points by handfuloflight on Aug 10, 2025 | hide | past | pdf | discuss