about
2431. SimpleQA Verified: Reliable Factuality Benchmark to Measure Parametric Knowledge (arxiv.org)
3 points by simonpure on Sep 10, 2025 | hide | past | pdf | discuss
2432. R-Zero: Self-Evolving Reasoning LLM from Zero Data (arxiv.org)
121 points by lawrenceyan on Sep 10, 2025 | hide | past | pdf | 61 comments
2433. Bootstrapping Task Spaces for Self-Improvement (arxiv.org)
2 points by Anon84 on Sep 9, 2025 | hide | past | pdf | discuss
2434. Maestro: Joint Graph and Config Optimization for Reliable AI Agents (arxiv.org)
2 points by knrz on Sep 9, 2025 | hide | past | pdf | discuss
2435. An AI system to help scientists write expert-level empirical software (arxiv.org)
3 points by Rudybega on Sep 9, 2025 | hide | past | pdf | 1 comment
2436. Outcome-Based Exploration for LLM Reasoning (arxiv.org)
2 points by badmonster on Sep 9, 2025 | hide | past | pdf | discuss
2437. An AI system to help scientists write expert-level empirical software (arxiv.org)
8 points by simonpure on Sep 9, 2025 | hide | past | pdf | 1 comment
2438. Curia: A Multi-Modal Foundation Model for Radiology (arxiv.org)
9 points by paulglx on Sep 9, 2025 | hide | past | pdf | discuss
2439. Set Block Decoding Is a Language Model Inference Accelerator (arxiv.org)
4 points by veryluckyxyz on Sep 9, 2025 | hide | past | pdf | discuss
2440. KVComp: A High-Performance, LLM-Aware, Lossy Compression Framework for KV Cache (arxiv.org)
3 points by kstonekuan on Sep 8, 2025 | hide | past | pdf | discuss
2441. Why Language Models Hallucinate (arxiv.org)
3 points by sonabinu on Sep 8, 2025 | hide | past | pdf | 1 comment
2442. A Comprehensive Survey on Trustworthiness in Reasoning with LLMs (arxiv.org)
1 point by Anon84 on Sep 8, 2025 | hide | past | pdf | discuss
2443. The LLM Has Left the Chat: Evidence of Bail Preferences in LLMs (arxiv.org)
1 point by geox on Sep 8, 2025 | hide | past | pdf | discuss
2444. Geometric Deep Learning Grids, Groups, Graphs, Geodesics, and Gauges [pdf] (arxiv.org)
3 points by ibobev on Sep 8, 2025 | hide | past | pdf | discuss
2445. How to Hack Transformers: Steering LLMs via Prompts, States, and Weight Edits (arxiv.org)
2 points by WASDAai on Sep 8, 2025 | hide | past | pdf | 1 comment
2446. Refrag: Rethinking RAG Based Decoding (arxiv.org)
4 points by itchyjunk on Sep 8, 2025 | hide | past | pdf | discuss
2447. PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning (arxiv.org)
1 point by PaulHoule on Sep 7, 2025 | hide | past | pdf | discuss
2448. Tricking LLM-Based NPCs into Spilling Secrets (arxiv.org)
4 points by PaulHoule on Sep 7, 2025 | hide | past | pdf | 2 comments
2449. Zero-Shot Reinforcement Learning (arxiv.org)
2 points by enjeeneer on Sep 7, 2025 | hide | past | pdf | discuss
2450. Contemplative Artificial Intelligence (arxiv.org)
3 points by lucasluitjes on Sep 6, 2025 | hide | past | pdf | 2 comments
2451. Do Language Models Agree with Human Perceptions of Suspense in Stories? (arxiv.org)
1 point by PaulHoule on Sep 5, 2025 | hide | past | pdf | discuss
2452. Fantastic pretraining optimizers and where to find them (arxiv.org)
42 points by fzliu on Sep 5, 2025 | hide | past | pdf | 4 comments
2453. Fleeting memory improves language learning but impairs reading time prediction (arxiv.org)
3 points by PaulHoule on Sep 5, 2025 | hide | past | pdf | discuss
2454. The Landscape of Agentic Reinforcement Learning for LLMs (arxiv.org)
4 points by sonabinu on Sep 4, 2025 | hide | past | pdf | discuss
2455. Survey of Deep Learning and Foundation Models for Time Series Forecasting (arxiv.org)
1 point by brandonb on Sep 4, 2025 | hide | past | pdf | discuss
2456. LLM Social Simulations Are a Promising Research Method (arxiv.org)
3 points by PaulHoule on Sep 4, 2025 | hide | past | pdf | discuss
2457. DaCe AD: Unifying High-Performance Automatic Differentiation for ML and SciComp (arxiv.org)
3 points by matt_d on Sep 4, 2025 | hide | past | pdf | discuss
2458. Generative Recommendations with Context Parallelism on Hierarchical Transducers (arxiv.org)
3 points by PaulHoule on Sep 3, 2025 | hide | past | pdf | discuss
2459. Monolingual speech recognition models beat multilingual models ~30x bigger (arxiv.org)
3 points by dbreunig on Sep 3, 2025 | hide | past | pdf | discuss
2460. Latent Theory of Mind: Decentralized Diffusion Arch for Co-Op Manipulation[pdf] (arxiv.org)
3 points by kelseyfrog on Sep 3, 2025 | hide | past | pdf | discuss