| 2431. |
SimpleQA Verified: Reliable Factuality Benchmark to Measure Parametric Knowledge (arxiv.org) |
|
3 points by simonpure on Sep 10, 2025 | hide | past | pdf | discuss
|
| 2432. |
R-Zero: Self-Evolving Reasoning LLM from Zero Data (arxiv.org) |
|
121 points by lawrenceyan on Sep 10, 2025 | hide | past | pdf | 61 comments
|
| 2433. |
Bootstrapping Task Spaces for Self-Improvement (arxiv.org) |
|
2 points by Anon84 on Sep 9, 2025 | hide | past | pdf | discuss
|
| 2434. |
Maestro: Joint Graph and Config Optimization for Reliable AI Agents (arxiv.org) |
|
2 points by knrz on Sep 9, 2025 | hide | past | pdf | discuss
|
| 2435. |
An AI system to help scientists write expert-level empirical software (arxiv.org) |
|
3 points by Rudybega on Sep 9, 2025 | hide | past | pdf | 1 comment
|
| 2436. |
Outcome-Based Exploration for LLM Reasoning (arxiv.org) |
|
2 points by badmonster on Sep 9, 2025 | hide | past | pdf | discuss
|
| 2437. |
An AI system to help scientists write expert-level empirical software (arxiv.org) |
|
8 points by simonpure on Sep 9, 2025 | hide | past | pdf | 1 comment
|
| 2438. |
Curia: A Multi-Modal Foundation Model for Radiology (arxiv.org) |
|
9 points by paulglx on Sep 9, 2025 | hide | past | pdf | discuss
|
| 2439. |
Set Block Decoding Is a Language Model Inference Accelerator (arxiv.org) |
|
4 points by veryluckyxyz on Sep 9, 2025 | hide | past | pdf | discuss
|
| 2440. |
KVComp: A High-Performance, LLM-Aware, Lossy Compression Framework for KV Cache (arxiv.org) |
|
3 points by kstonekuan on Sep 8, 2025 | hide | past | pdf | discuss
|
| 2441. |
Why Language Models Hallucinate (arxiv.org) |
|
3 points by sonabinu on Sep 8, 2025 | hide | past | pdf | 1 comment
|
| 2442. |
A Comprehensive Survey on Trustworthiness in Reasoning with LLMs (arxiv.org) |
|
1 point by Anon84 on Sep 8, 2025 | hide | past | pdf | discuss
|
| 2443. |
The LLM Has Left the Chat: Evidence of Bail Preferences in LLMs (arxiv.org) |
|
1 point by geox on Sep 8, 2025 | hide | past | pdf | discuss
|
| 2444. |
Geometric Deep Learning Grids, Groups, Graphs, Geodesics, and Gauges [pdf] (arxiv.org) |
|
3 points by ibobev on Sep 8, 2025 | hide | past | pdf | discuss
|
| 2445. |
How to Hack Transformers: Steering LLMs via Prompts, States, and Weight Edits (arxiv.org) |
|
2 points by WASDAai on Sep 8, 2025 | hide | past | pdf | 1 comment
|
| 2446. |
Refrag: Rethinking RAG Based Decoding (arxiv.org) |
|
4 points by itchyjunk on Sep 8, 2025 | hide | past | pdf | discuss
|
| 2447. |
PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning (arxiv.org) |
|
1 point by PaulHoule on Sep 7, 2025 | hide | past | pdf | discuss
|
| 2448. |
Tricking LLM-Based NPCs into Spilling Secrets (arxiv.org) |
|
4 points by PaulHoule on Sep 7, 2025 | hide | past | pdf | 2 comments
|
| 2449. |
Zero-Shot Reinforcement Learning (arxiv.org) |
|
2 points by enjeeneer on Sep 7, 2025 | hide | past | pdf | discuss
|
| 2450. |
Contemplative Artificial Intelligence (arxiv.org) |
|
3 points by lucasluitjes on Sep 6, 2025 | hide | past | pdf | 2 comments
|
| 2451. |
Do Language Models Agree with Human Perceptions of Suspense in Stories? (arxiv.org) |
|
1 point by PaulHoule on Sep 5, 2025 | hide | past | pdf | discuss
|
| 2452. |
Fantastic pretraining optimizers and where to find them (arxiv.org) |
|
42 points by fzliu on Sep 5, 2025 | hide | past | pdf | 4 comments
|
| 2453. |
Fleeting memory improves language learning but impairs reading time prediction (arxiv.org) |
|
3 points by PaulHoule on Sep 5, 2025 | hide | past | pdf | discuss
|
| 2454. |
The Landscape of Agentic Reinforcement Learning for LLMs (arxiv.org) |
|
4 points by sonabinu on Sep 4, 2025 | hide | past | pdf | discuss
|
| 2455. |
Survey of Deep Learning and Foundation Models for Time Series Forecasting (arxiv.org) |
|
1 point by brandonb on Sep 4, 2025 | hide | past | pdf | discuss
|
| 2456. |
LLM Social Simulations Are a Promising Research Method (arxiv.org) |
|
3 points by PaulHoule on Sep 4, 2025 | hide | past | pdf | discuss
|
| 2457. |
DaCe AD: Unifying High-Performance Automatic Differentiation for ML and SciComp (arxiv.org) |
|
3 points by matt_d on Sep 4, 2025 | hide | past | pdf | discuss
|
| 2458. |
Generative Recommendations with Context Parallelism on Hierarchical Transducers (arxiv.org) |
|
3 points by PaulHoule on Sep 3, 2025 | hide | past | pdf | discuss
|
| 2459. |
Monolingual speech recognition models beat multilingual models ~30x bigger (arxiv.org) |
|
3 points by dbreunig on Sep 3, 2025 | hide | past | pdf | discuss
|
| 2460. |
Latent Theory of Mind: Decentralized Diffusion Arch for Co-Op Manipulation[pdf] (arxiv.org) |
|
3 points by kelseyfrog on Sep 3, 2025 | hide | past | pdf | discuss
|
| More |