| 3031. |
Mind the Gap: Deep Learning Doesn't Learn Deeply (arxiv.org) |
|
2 points by nyrikki on May 30, 2025 | hide | past | pdf | discuss
|
| 3032. |
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning (arxiv.org) |
|
2 points by Anon84 on May 30, 2025 | hide | past | pdf | discuss
|
| 3033. |
MuLoCo: Muon is a practical inner optimizer for DiLoCo (arxiv.org) |
|
2 points by Mougatine on May 30, 2025 | hide | past | pdf | discuss
|
| 3034. |
Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents (arxiv.org) |
|
7 points by hardmaru on May 30, 2025 | hide | past | pdf | 1 comment
|
| 3035. |
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence (arxiv.org) |
|
2 points by badmonster on May 30, 2025 | hide | past | pdf | 1 comment
|
| 3036. |
SWE-Rebench: Task Collection and Decontaminated Evaluation of SWE Agents (arxiv.org) |
|
2 points by SerCe on May 30, 2025 | hide | past | pdf | discuss
|
| 3037. |
Superhuman performance of an LLM on the reasoning tasks of a physician (arxiv.org) |
|
36 points by amichail on May 29, 2025 | hide | past | pdf | 31 comments
|
| 3038. |
A Practical Deep Learning-Based Acoustic Side Channel Attack on Keyboards (arxiv.org) |
|
2 points by bookofjoe on May 29, 2025 | hide | past | pdf | discuss
|
| 3039. |
Towards Large-Scale Generative Ranking (arxiv.org) |
|
2 points by PaulHoule on May 29, 2025 | hide | past | pdf | discuss
|
| 3040. |
Pre-Training for Recommendation Unlearning (arxiv.org) |
|
1 point by badmonster on May 29, 2025 | hide | past | pdf | discuss
|
| 3041. |
VideoGameBench: Can Vision-Language Models complete popular video games? (arxiv.org) |
|
4 points by wertyk on May 29, 2025 | hide | past | pdf | 1 comment
|
| 3042. |
Sufficient Context: A New Lens on Retrieval Augmented Generation Systems (arxiv.org) |
|
3 points by simonpure on May 29, 2025 | hide | past | pdf | discuss
|
| 3043. |
Science Board: Evaluating Agents in Realistic Scientific Workflows (arxiv.org) |
|
7 points by simonpure on May 29, 2025 | hide | past | pdf | discuss
|
| 3044. |
MMaDA: Multimodal Large Diffusion Language Models (arxiv.org) |
|
1 point by doener on May 28, 2025 | hide | past | pdf | discuss
|
| 3045. |
AR-Diffusion: Auto-Regressive Diffusion Model for Text Generation (arxiv.org) |
|
7 points by doener on May 28, 2025 | hide | past | pdf | 3 comments
|
| 3046. |
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective (arxiv.org) |
|
19 points by doener on May 28, 2025 | hide | past | pdf | 1 comment
|
| 3047. |
Can Large Reasoning Models Self-Train? (arxiv.org) |
|
1 point by belter on May 28, 2025 | hide | past | pdf | discuss
|
| 3048. |
When Models Don't Collapse: On the Consistency of Iterative MLE (arxiv.org) |
|
1 point by belter on May 28, 2025 | hide | past | pdf | discuss
|
| 3049. |
New Lens on RAG Systems (arxiv.org) |
|
1 point by omarsar on May 28, 2025 | hide | past | pdf | discuss
|
| 3050. |
FlowTSE: Target Speaker Extraction with Flow Matching (arxiv.org) |
|
25 points by agold97 on May 28, 2025 | hide | past | pdf | 2 comments
|
| 3051. |
Embedding-to-Prefix: Spotify's Efficient Personalization for LLMs (arxiv.org) |
|
1 point by willvarfar on May 28, 2025 | hide | past | pdf | discuss
|
| 3052. |
Sea-Helm: Southeast Asian Holistic Evaluation of Language Models (arxiv.org) |
|
2 points by JSR_FDED on May 28, 2025 | hide | past | pdf | discuss
|
| 3053. |
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents (arxiv.org) |
|
3 points by gfto on May 28, 2025 | hide | past | pdf | discuss
|
| 3054. |
Synthetic Data RL: Task Definition Is All You Need (arxiv.org) |
|
2 points by simonpure on May 28, 2025 | hide | past | pdf | discuss
|
| 3055. |
Optimization by unifying stochastic gradient and quasi-Newton methods (2013) (arxiv.org) |
|
3 points by fzliu on May 27, 2025 | hide | past | pdf | discuss
|
| 3056. |
Self-Reflective Uncertainties: Do LLMs Know Their Internal Answer Distribution? (arxiv.org) |
|
1 point by badmonster on May 27, 2025 | hide | past | pdf | discuss
|
| 3057. |
Arc-NCA: Towards Developmental Solutions to the Abstraction and Reasoning Corpus (arxiv.org) |
|
1 point by jarmitage on May 27, 2025 | hide | past | pdf | discuss
|
| 3058. |
Frontier Models are Capable of In-context Scheming (arxiv.org) |
|
1 point by doener on May 27, 2025 | hide | past | pdf | discuss
|
| 3059. |
Extracting memorized pieces of books from open-weight language models (arxiv.org) |
|
2 points by Tomte on May 27, 2025 | hide | past | pdf | discuss
|
| 3060. |
Deep Reinforcement Learning, a Textbook (2023) (arxiv.org) |
|
4 points by Anon84 on May 27, 2025 | hide | past | pdf | discuss
|
| More |