about
3361. Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad (arxiv.org)
2 points by ydnyshhh on Apr 8, 2025 | hide | past | pdf | discuss
3362. Rope to Nope and Back Again: A New Hybrid Attention Strategy (arxiv.org)
3 points by ydnyshhh on Apr 8, 2025 | hide | past | pdf | discuss
3363. InfiniteICL: Breaking the Limit of Context Window Size (arxiv.org)
1 point by demirbey05 on Apr 8, 2025 | hide | past | pdf | discuss
3364. Transformers Are Efficient Compilers, Provably (arxiv.org)
5 points by wseqyrku on Apr 7, 2025 | hide | past | pdf | discuss
3365. How Does Watermarking Affect Visual Language Models in Document Understanding? (arxiv.org)
1 point by PaulHoule on Apr 7, 2025 | hide | past | pdf | discuss
3366. Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad (arxiv.org)
3 points by marvinborner on Apr 7, 2025 | hide | past | pdf | discuss
3367. Thought2Text: Text Generation from EEG Signal Using Large Language Models (arxiv.org)
2 points by trkaky on Apr 7, 2025 | hide | past | pdf | discuss
3368. SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random Generators (arxiv.org)
1 point by Anon84 on Apr 7, 2025 | hide | past | pdf | discuss
3369. UniGen: Unified Modeling of Initial Agent States and Trajectories (arxiv.org)
1 point by lawrenceyan on Apr 6, 2025 | hide | past | pdf | discuss
3370. ScienceWorld: Is your Agent Smarter than a 5th Grader? (arxiv.org)
2 points by liamdgray on Apr 6, 2025 | hide | past | pdf | 1 comment
3371. Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad (arxiv.org)
2 points by rahimnathwani on Apr 6, 2025 | hide | past | pdf | discuss
3372. Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad (arxiv.org)
5 points by Bluestein on Apr 6, 2025 | hide | past | pdf | discuss
3373. Advances and Challenges in Foundation Agents (arxiv.org)
1 point by Anon84 on Apr 5, 2025 | hide | past | pdf | discuss
3374. Faith and Fate: Limits of Transformers on Compositionality (arxiv.org)
3 points by gsf_emergency_2 on Apr 5, 2025 | hide | past | pdf | 1 comment
3375. Generating Medically-Informed Explanations for Depression Detection Using LLMs (arxiv.org)
1 point by PaulHoule on Apr 4, 2025 | hide | past | pdf | discuss
3376. Scaling Language-Free Visual Representation Learning (arxiv.org)
3 points by Anon84 on Apr 4, 2025 | hide | past | pdf | discuss
3377. Fine-Tuning LLMs for Report Summarization (arxiv.org)
1 point by PaulHoule on Apr 4, 2025 | hide | past | pdf | discuss
3378. DeepSeek: Inference-Time Scaling for Generalist Reward Modeling (arxiv.org)
163 points by tim_sw on Apr 4, 2025 | hide | past | pdf | 35 comments
3379. Measuring AI Ability to Complete Long Tasks (arxiv.org)
2 points by chriskanan on Apr 3, 2025 | hide | past | pdf | discuss
3380. Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad (arxiv.org)
2 points by mromanuk on Apr 3, 2025 | hide | past | pdf | discuss
3381. Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers (arxiv.org)
2 points by geox on Apr 3, 2025 | hide | past | pdf | discuss
3382. Search-R1: Training LLMs to Reason and Leverage Search Engines with RL (arxiv.org)
101 points by jonbaer on Apr 3, 2025 | hide | past | pdf | 12 comments
3383. Cordic Is All You Need (arxiv.org)
2 points by PaulHoule on Apr 2, 2025 | hide | past | pdf | discuss
3384. Perceiver IO: A General Architecture for Structured Inputs and Outputs (2021) (arxiv.org)
2 points by Topfi on Apr 2, 2025 | hide | past | pdf | discuss
3385. Evaluating Agent-Based Program Repair at Google (arxiv.org)
15 points by azhenley on Apr 2, 2025 | hide | past | pdf | 1 comment
3386. Multi-Token Attention (arxiv.org)
152 points by fzliu on Apr 2, 2025 | hide | past | pdf | 44 comments
3387. Human-in-the-Loop Local Corrections of 3D Scene Layouts via Infilling (arxiv.org)
2 points by PaulHoule on Apr 2, 2025 | hide | past | pdf | discuss
3388. Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters (arxiv.org)
1 point by consumer451 on Apr 2, 2025 | hide | past | pdf | 1 comment
3389. HiRAG: RAG with Hierarchical Knowledge (arxiv.org)
3 points by fmos on Apr 2, 2025 | hide | past | pdf | 1 comment
3390. UCSD: Large Language Models Pass the Turing Test (arxiv.org)
91 points by Mossy9 on Apr 2, 2025 | hide | past | pdf | 106 comments