| 3361. |
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad (arxiv.org) |
|
2 points by ydnyshhh on Apr 8, 2025 | hide | past | pdf | discuss
|
| 3362. |
Rope to Nope and Back Again: A New Hybrid Attention Strategy (arxiv.org) |
|
3 points by ydnyshhh on Apr 8, 2025 | hide | past | pdf | discuss
|
| 3363. |
InfiniteICL: Breaking the Limit of Context Window Size (arxiv.org) |
|
1 point by demirbey05 on Apr 8, 2025 | hide | past | pdf | discuss
|
| 3364. |
Transformers Are Efficient Compilers, Provably (arxiv.org) |
|
5 points by wseqyrku on Apr 7, 2025 | hide | past | pdf | discuss
|
| 3365. |
How Does Watermarking Affect Visual Language Models in Document Understanding? (arxiv.org) |
|
1 point by PaulHoule on Apr 7, 2025 | hide | past | pdf | discuss
|
| 3366. |
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad (arxiv.org) |
|
3 points by marvinborner on Apr 7, 2025 | hide | past | pdf | discuss
|
| 3367. |
Thought2Text: Text Generation from EEG Signal Using Large Language Models (arxiv.org) |
|
2 points by trkaky on Apr 7, 2025 | hide | past | pdf | discuss
|
| 3368. |
SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random Generators (arxiv.org) |
|
1 point by Anon84 on Apr 7, 2025 | hide | past | pdf | discuss
|
| 3369. |
UniGen: Unified Modeling of Initial Agent States and Trajectories (arxiv.org) |
|
1 point by lawrenceyan on Apr 6, 2025 | hide | past | pdf | discuss
|
| 3370. |
ScienceWorld: Is your Agent Smarter than a 5th Grader? (arxiv.org) |
|
2 points by liamdgray on Apr 6, 2025 | hide | past | pdf | 1 comment
|
| 3371. |
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad (arxiv.org) |
|
2 points by rahimnathwani on Apr 6, 2025 | hide | past | pdf | discuss
|
| 3372. |
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad (arxiv.org) |
|
5 points by Bluestein on Apr 6, 2025 | hide | past | pdf | discuss
|
| 3373. |
Advances and Challenges in Foundation Agents (arxiv.org) |
|
1 point by Anon84 on Apr 5, 2025 | hide | past | pdf | discuss
|
| 3374. |
Faith and Fate: Limits of Transformers on Compositionality (arxiv.org) |
|
3 points by gsf_emergency_2 on Apr 5, 2025 | hide | past | pdf | 1 comment
|
| 3375. |
Generating Medically-Informed Explanations for Depression Detection Using LLMs (arxiv.org) |
|
1 point by PaulHoule on Apr 4, 2025 | hide | past | pdf | discuss
|
| 3376. |
Scaling Language-Free Visual Representation Learning (arxiv.org) |
|
3 points by Anon84 on Apr 4, 2025 | hide | past | pdf | discuss
|
| 3377. |
Fine-Tuning LLMs for Report Summarization (arxiv.org) |
|
1 point by PaulHoule on Apr 4, 2025 | hide | past | pdf | discuss
|
| 3378. |
DeepSeek: Inference-Time Scaling for Generalist Reward Modeling (arxiv.org) |
|
163 points by tim_sw on Apr 4, 2025 | hide | past | pdf | 35 comments
|
| 3379. |
Measuring AI Ability to Complete Long Tasks (arxiv.org) |
|
2 points by chriskanan on Apr 3, 2025 | hide | past | pdf | discuss
|
| 3380. |
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad (arxiv.org) |
|
2 points by mromanuk on Apr 3, 2025 | hide | past | pdf | discuss
|
| 3381. |
Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers (arxiv.org) |
|
2 points by geox on Apr 3, 2025 | hide | past | pdf | discuss
|
| 3382. |
Search-R1: Training LLMs to Reason and Leverage Search Engines with RL (arxiv.org) |
|
101 points by jonbaer on Apr 3, 2025 | hide | past | pdf | 12 comments
|
| 3383. |
Cordic Is All You Need (arxiv.org) |
|
2 points by PaulHoule on Apr 2, 2025 | hide | past | pdf | discuss
|
| 3384. |
Perceiver IO: A General Architecture for Structured Inputs and Outputs (2021) (arxiv.org) |
|
2 points by Topfi on Apr 2, 2025 | hide | past | pdf | discuss
|
| 3385. |
Evaluating Agent-Based Program Repair at Google (arxiv.org) |
|
15 points by azhenley on Apr 2, 2025 | hide | past | pdf | 1 comment
|
| 3386. |
Multi-Token Attention (arxiv.org) |
|
152 points by fzliu on Apr 2, 2025 | hide | past | pdf | 44 comments
|
| 3387. |
Human-in-the-Loop Local Corrections of 3D Scene Layouts via Infilling (arxiv.org) |
|
2 points by PaulHoule on Apr 2, 2025 | hide | past | pdf | discuss
|
| 3388. |
Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters (arxiv.org) |
|
1 point by consumer451 on Apr 2, 2025 | hide | past | pdf | 1 comment
|
| 3389. |
HiRAG: RAG with Hierarchical Knowledge (arxiv.org) |
|
3 points by fmos on Apr 2, 2025 | hide | past | pdf | 1 comment
|
| 3390. |
UCSD: Large Language Models Pass the Turing Test (arxiv.org) |
|
91 points by Mossy9 on Apr 2, 2025 | hide | past | pdf | 106 comments
|
| More |